A fast file deduplicator
-
Updated
Oct 10, 2025 - Rust
A fast file deduplicator
Yet Another Dupes Finder
Smash through to find duplicate files super fast by slicing files intelligently!
.NET 10 API for document file format identification, text/metadata/attachment/embedded object/sensitive item (PII/PHI/FERPA)/entity extraction.
Case study using dotfurther's Open Discover Platform with the RavenDB document store to rapidly create a full-text search/eDiscovery/information governance capable demonstration application.
fast, parallel duplicate file detection
A safe, fast, and controllable duplicate file manager for the command line. Not just a duplicate finder. It is a complete workflow for locating, reviewing, and cleaning duplicate files — with byte-level accuracy, reversible actions, multi-stage detection, and full user control at every step.
File deduplicator
a small C++ CLI for finding duplicate files (namely, a deduplicator..) using size, hashing, and byte-by-byte comparison!! this really helped me find, like, 23 files (turned out to be the same powerpoint presentation that i spammed because it wasn't showing in the downloads..)
Duplicate file finder and de-duplicator. A tool that detects duplicate files and replaces them with symlinks to a shared file in a special shared files folder.
A simple command line utility that identifies groups of identical files and displays them to the console.
CloneZapper is a Python script that hunts down identical files within a directory and its subdirectories, ruthlessly eliminating them!
Этот проект представляет собой мощный инструмент для поиска и анализа дублирующихся файлов в указанной директории. Программа позволяет эффективно выявлять одинаковые файлы на основе их содержимого, используя алгоритм хеширования SHA-256. Она поддерживает настройку параметров, таких как минимальный размер файла для проверки и игнорирование определен
Production-oriented file deduplication and data-protection platform with tenant isolation, SHA-256 exact deduplication, MinHash/LSH and dHash near-duplicate detection, DLP, AES-256-GCM encryption, RBAC, and audit logging.
A Node.js CLI tool that recursively scans directories, detects duplicate files by content using hashing, and safely removes redundant copies with dry-run support.
A corpus-hygiene utility for RAG data pipelines that identifies duplicate content risk, quantifies duplication with actionable statistics, and supports controlled remediation before indexing. It enables staged audit-then-cull workflows that improve retrieval quality, reduce embedding/indexing cost, and strengthen governance in knowledge curation.
Golang Deduplication Utility Library
Find duplicate files in multiple folder(s) scanning .txt or/and .torrent files and depending on the selected mode (readonly: true | false) get information about duplicated files /+ extract them into new folders
To associate your repository with the file-deduplication topic, visit your repo's landing page and select "manage topics."