A fast, parallel Rust CLI that finds and quarantines duplicate and visually-similar images using a multi-stage filtering pipeline. Elapsed time is printed after every run.
Detection runs in five phases:
Recursively walks the target directory (skipping any existing duplicates/ subfolder) and collects every supported image file. Each file's metadata (path, size, creation time) is read in parallel via rayon. Files that cannot be stat'd are silently skipped.
Supported formats: .jpg / .jpeg, .png, .webp, .bmp, .tiff
Files are grouped by byte size (a fast O(1) pre-filter). Within each size group, a SHA-256 hash is computed in parallel. Files with identical hashes are byte-for-byte duplicates.
Files already marked for quarantine by Phase 2 are excluded; every other image — including the keeper of each exact-duplicate group — is hashed with a 16×16 dHash (difference hash), so a keeper can still match images that are visually similar without being byte-identical:
- Each image is resized to 17×16 pixels with a triangle (bilinear) filter.
- The grayscale thumbnail is converted to a 256-bit integer by comparing each pixel to its right neighbour across 16 rows × 16 columns = 256 comparisons, stored as
[u64; 4]. - Pairwise comparison is fully parallelised: each row
iof the hash matrix is processed by a separate rayon thread, scanning rowsi+1..nfor matches —O(n²)work spread across all CPU cores. - Pairs whose hashes differ by ≤ 40 bits (Hamming distance, ~15.6% of 256 bits) are connected.
- A sequential union-find (with path-halving compression) merges transitive similarity chains into correct groups (e.g. A≈B and B≈C → single group {A, B, C}).
A threshold of 40/256 bits catches resized, lightly re-compressed, or slightly-edited versions of the same image while avoiding false positives between genuinely different images.
Within each duplicate group (exact or perceptual), the oldest file (by creation time, falling back to modification time on platforms that don't record birth time) is kept in place. All newer files are marked for quarantine.
Log records are built before any files are moved so scores are computed from the original paths. Duplicates are then moved into <target-dir>/duplicates/, preserving their original subdirectory structure to prevent filename collisions. If a destination path is already taken (typically from an earlier run), _2, _3, … is inserted before the extension rather than overwriting it. A duplicate_log.csv is written to <target-dir>/, next to the duplicates/ folder.
target-dir/
├── oldest_original.jpg ← kept
├── another_unique.png ← kept
├── duplicate_log.csv ← written here
└── duplicates/
├── newer_copy.jpg ← moved here
└── resized_version.jpeg ← moved here
| Column | Description |
|---|---|
original_path |
Path of the kept (oldest) image |
duplicate_path |
Path the duplicate was moved from |
match_type |
exact_hash or perceptual_similarity |
similarity_score |
0 for exact matches; Hamming distance in bits (out of 256) for perceptual matches |
kept_version_reason |
Always oldest_file |
image-dedup [OPTIONS] <DIR>
Pause syncing before running on a cloud-synced folder. If the target directory lives on Google Drive, OneDrive, Dropbox, iCloud Drive, or any similar service, pause or disable syncing for the duration of the run and let it re-sync afterwards. Moving a large number of files at once fights the sync client: it may re-download files as they are moved out, generate "conflicted copy" duplicates that then look like new duplicates on the next run, hold locks that make moves fail partway through, or propagate the moves to your other machines before you have reviewed the results. Placeholder/online-only files (OneDrive Files On-Demand, Drive streaming mode) are a further problem — reading them forces a download of every image in the tree.
Run
--dry-runfirst regardless; it touches no files.
| Argument | Description |
|---|---|
<DIR> |
Directory to scan recursively for images |
| Flag | Description |
|---|---|
-f, --force |
Skip the confirmation prompt and proceed immediately |
--dry-run |
Analyse and print what would be moved; touch no files |
-h, --help |
Print help |
# Preview what would be moved (safe, no changes)
image-dedup --dry-run ~/Pictures
# Run interactively (prompts before moving)
image-dedup ~/Pictures
# Run non-interactively in scripts
image-dedup --force ~/Pictures
# Combine dry-run with a nested folder
image-dedup --dry-run ./photos/vacation/2024cargo install --git https://github.com/BluePlexus/image-dedupThis builds in release mode and places image-dedup on your PATH (in ~/.cargo/bin). Requires a Rust toolchain — see rustup.rs.
To build from a local checkout instead:
cargo install --path .# Debug build (development / testing)
cargo build
# Release build — ~9× faster, recommended for large libraries
cargo build --release
# Binaries at:
# target/debug/image-dedup
# target/release/image-dedupcargo testEverything runs offline in temp directories; no fixture images are checked in (they are generated deterministically at test time).
| Suite | Location | What it covers |
|---|---|---|
| Unit | src/lib.rs mod tests |
dHash, Hamming distance, threshold boundaries, keeper selection, union-find clustering, log records |
| Integration | tests/cli.rs | The real binary end-to-end: file movement, CSV contents, dry-run, re-run idempotency, collision suffixing, exit codes |
Run a single test by name:
cargo test keeper_is_the_oldestCI (.github/workflows/ci.yml) runs cargo fmt --check,
cargo clippy -D warnings, and the full suite on Linux and macOS for every push
and pull request. Tests run on both because creation-time behaviour differs
between platforms and keeper selection depends on it.
| Crate | Purpose |
|---|---|
image 0.24 |
Image decoding, resizing, grayscale conversion |
sha2 0.10 |
SHA-256 for exact duplicate hashing |
rayon 1.8 |
Data-parallel iterators (metadata, hashing, O(n²) comparison) |
walkdir 2.4 |
Recursive directory traversal |
clap 4.4 |
CLI argument parsing |
indicatif 0.17 |
Progress bars |
dialoguer 0.11 |
Interactive confirmation prompt |
csv 1.3 |
CSV log output |
anyhow 1.0 |
Ergonomic error handling |
16×16 dHash over 8×8
The hash was upgraded from 8×8 (64-bit, u64) to 16×16 (256-bit, [u64; 4]). The larger hash has finer per-pixel granularity and catches fewer false positives between images that are similar but not duplicates. The Hamming threshold scales proportionally: 40/256 ≈ 15.6%, matching the former 10/64 ratio.
Oldest file kept The keeper criterion changed from highest resolution/largest filesize to oldest creation timestamp. Creation time is a better proxy for "original" in photo library deduplication — the original capture is typically the oldest copy, while rescaled exports, web downloads, or backup copies arrive later. On filesystems that do not record birth time (some Linux ext4 mounts), modification time is used as a fallback.
Parallel O(n²) comparison with union-find
The perceptual comparison loop is now parallelised: rayon's par_iter().flat_map_iter() assigns each image index i to a thread that sequentially walks indices i+1..n — the only safe decomposition for a triangular work matrix. Results are collected as (i, j) pairs and fed into a sequential union-find, which correctly handles transitive chains (A≈B, B≈C → keep oldest of {A, B, C}) and avoids the double-assignment bugs that come from simple O(n²) loops with a used set.
Log records computed before file moves
Perceptual similarity scores (Hamming distances) are computed and stored during group construction, then written to log_records before any fs::rename calls. This avoids the earlier design's bug where the score was recomputed by re-opening an already-moved file.
Phase 1 no longer decodes images Metadata extraction (Phase 1) now reads only filesystem metadata — no image decoding. This makes Phase 1 nearly instant. Invalid or corrupt files that pass the extension check are simply skipped later in the SHA-256 (Phase 2) or dHash (Phase 3) phases, where decoding happens anyway.
Re-run safety
The duplicates/ folder is excluded from scanning on every run, so re-running on the same directory will not re-process already-quarantined files.
MIT — see LICENSE.