Skip to content

Latest commit

 

History

11 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

image-dedup — Image Duplicate Detector

A fast, parallel Rust CLI that finds and quarantines duplicate and visually-similar images using a multi-stage filtering pipeline. Elapsed time is printed after every run.


How it works

Detection runs in five phases:

Phase 1 — Scan & metadata extraction

Recursively walks the target directory (skipping any existing duplicates/ subfolder) and collects every supported image file. Each file's metadata (path, size, creation time) is read in parallel via rayon. Files that cannot be stat'd are silently skipped.

Supported formats: .jpg / .jpeg, .png, .webp, .bmp, .tiff

Phase 2 — Exact duplicate detection

Files are grouped by byte size (a fast O(1) pre-filter). Within each size group, a SHA-256 hash is computed in parallel. Files with identical hashes are byte-for-byte duplicates.

Phase 3 — Perceptual similarity detection (parallel)

Files already marked for quarantine by Phase 2 are excluded; every other image — including the keeper of each exact-duplicate group — is hashed with a 16×16 dHash (difference hash), so a keeper can still match images that are visually similar without being byte-identical:

  1. Each image is resized to 17×16 pixels with a triangle (bilinear) filter.
  2. The grayscale thumbnail is converted to a 256-bit integer by comparing each pixel to its right neighbour across 16 rows × 16 columns = 256 comparisons, stored as [u64; 4].
  3. Pairwise comparison is fully parallelised: each row i of the hash matrix is processed by a separate rayon thread, scanning rows i+1..n for matches — O(n²) work spread across all CPU cores.
  4. Pairs whose hashes differ by ≤ 40 bits (Hamming distance, ~15.6% of 256 bits) are connected.
  5. A sequential union-find (with path-halving compression) merges transitive similarity chains into correct groups (e.g. A≈B and B≈C → single group {A, B, C}).

A threshold of 40/256 bits catches resized, lightly re-compressed, or slightly-edited versions of the same image while avoiding false positives between genuinely different images.

Phase 4 — Keeper selection

Within each duplicate group (exact or perceptual), the oldest file (by creation time, falling back to modification time on platforms that don't record birth time) is kept in place. All newer files are marked for quarantine.

Phase 5 — File movement & logging

Log records are built before any files are moved so scores are computed from the original paths. Duplicates are then moved into <target-dir>/duplicates/, preserving their original subdirectory structure to prevent filename collisions. If a destination path is already taken (typically from an earlier run), _2, _3, … is inserted before the extension rather than overwriting it. A duplicate_log.csv is written to <target-dir>/, next to the duplicates/ folder.


Output structure

target-dir/
├── oldest_original.jpg         ← kept
├── another_unique.png          ← kept
├── duplicate_log.csv           ← written here
└── duplicates/
    ├── newer_copy.jpg          ← moved here
    └── resized_version.jpeg    ← moved here

duplicate_log.csv columns

Column Description
original_path Path of the kept (oldest) image
duplicate_path Path the duplicate was moved from
match_type exact_hash or perceptual_similarity
similarity_score 0 for exact matches; Hamming distance in bits (out of 256) for perceptual matches
kept_version_reason Always oldest_file

Usage

image-dedup [OPTIONS] <DIR>

Pause syncing before running on a cloud-synced folder. If the target directory lives on Google Drive, OneDrive, Dropbox, iCloud Drive, or any similar service, pause or disable syncing for the duration of the run and let it re-sync afterwards. Moving a large number of files at once fights the sync client: it may re-download files as they are moved out, generate "conflicted copy" duplicates that then look like new duplicates on the next run, hold locks that make moves fail partway through, or propagate the moves to your other machines before you have reviewed the results. Placeholder/online-only files (OneDrive Files On-Demand, Drive streaming mode) are a further problem — reading them forces a download of every image in the tree.

Run --dry-run first regardless; it touches no files.

Arguments

Argument Description
<DIR> Directory to scan recursively for images

Options

Flag Description
-f, --force Skip the confirmation prompt and proceed immediately
--dry-run Analyse and print what would be moved; touch no files
-h, --help Print help

Examples

# Preview what would be moved (safe, no changes)
image-dedup --dry-run ~/Pictures

# Run interactively (prompts before moving)
image-dedup ~/Pictures

# Run non-interactively in scripts
image-dedup --force ~/Pictures

# Combine dry-run with a nested folder
image-dedup --dry-run ./photos/vacation/2024

Install

cargo install --git https://github.com/BluePlexus/image-dedup

This builds in release mode and places image-dedup on your PATH (in ~/.cargo/bin). Requires a Rust toolchain — see rustup.rs.

To build from a local checkout instead:

cargo install --path .

Build

# Debug build (development / testing)
cargo build

# Release build — ~9× faster, recommended for large libraries
cargo build --release

# Binaries at:
#   target/debug/image-dedup
#   target/release/image-dedup

Tests

cargo test

Everything runs offline in temp directories; no fixture images are checked in (they are generated deterministically at test time).

Suite Location What it covers
Unit src/lib.rs mod tests dHash, Hamming distance, threshold boundaries, keeper selection, union-find clustering, log records
Integration tests/cli.rs The real binary end-to-end: file movement, CSV contents, dry-run, re-run idempotency, collision suffixing, exit codes

Run a single test by name:

cargo test keeper_is_the_oldest

CI (.github/workflows/ci.yml) runs cargo fmt --check, cargo clippy -D warnings, and the full suite on Linux and macOS for every push and pull request. Tests run on both because creation-time behaviour differs between platforms and keeper selection depends on it.


Dependencies

Crate Purpose
image 0.24 Image decoding, resizing, grayscale conversion
sha2 0.10 SHA-256 for exact duplicate hashing
rayon 1.8 Data-parallel iterators (metadata, hashing, O(n²) comparison)
walkdir 2.4 Recursive directory traversal
clap 4.4 CLI argument parsing
indicatif 0.17 Progress bars
dialoguer 0.11 Interactive confirmation prompt
csv 1.3 CSV log output
anyhow 1.0 Ergonomic error handling

Design notes & trade-offs

16×16 dHash over 8×8 The hash was upgraded from 8×8 (64-bit, u64) to 16×16 (256-bit, [u64; 4]). The larger hash has finer per-pixel granularity and catches fewer false positives between images that are similar but not duplicates. The Hamming threshold scales proportionally: 40/256 ≈ 15.6%, matching the former 10/64 ratio.

Oldest file kept The keeper criterion changed from highest resolution/largest filesize to oldest creation timestamp. Creation time is a better proxy for "original" in photo library deduplication — the original capture is typically the oldest copy, while rescaled exports, web downloads, or backup copies arrive later. On filesystems that do not record birth time (some Linux ext4 mounts), modification time is used as a fallback.

Parallel O(n²) comparison with union-find The perceptual comparison loop is now parallelised: rayon's par_iter().flat_map_iter() assigns each image index i to a thread that sequentially walks indices i+1..n — the only safe decomposition for a triangular work matrix. Results are collected as (i, j) pairs and fed into a sequential union-find, which correctly handles transitive chains (A≈B, B≈C → keep oldest of {A, B, C}) and avoids the double-assignment bugs that come from simple O(n²) loops with a used set.

Log records computed before file moves Perceptual similarity scores (Hamming distances) are computed and stored during group construction, then written to log_records before any fs::rename calls. This avoids the earlier design's bug where the score was recomputed by re-opening an already-moved file.

Phase 1 no longer decodes images Metadata extraction (Phase 1) now reads only filesystem metadata — no image decoding. This makes Phase 1 nearly instant. Invalid or corrupt files that pass the extension check are simply skipped later in the SHA-256 (Phase 2) or dHash (Phase 3) phases, where decoding happens anyway.

Re-run safety The duplicates/ folder is excluded from scanning on every run, so re-running on the same directory will not re-process already-quarantined files.


License

MIT — see LICENSE.

About

Fast parallel Rust CLI for finding exact and visually-similar duplicate images

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages