Skip to content
#

file-deduplication

Here are 35 public repositories matching this topic...

Case study using dotfurther's Open Discover Platform with the RavenDB document store to rapidly create a full-text search/eDiscovery/information governance capable demonstration application.

  • Updated May 28, 2024

A safe, fast, and controllable duplicate file manager for the command line. Not just a duplicate finder. It is a complete workflow for locating, reviewing, and cleaning duplicate files — with byte-level accuracy, reversible actions, multi-stage detection, and full user control at every step.

  • Updated Sep 30, 2026
  • Rust

Этот проект представляет собой мощный инструмент для поиска и анализа дублирующихся файлов в указанной директории. Программа позволяет эффективно выявлять одинаковые файлы на основе их содержимого, используя алгоритм хеширования SHA-256. Она поддерживает настройку параметров, таких как минимальный размер файла для проверки и игнорирование определен

  • Updated Feb 14, 2025
  • Python

Production-oriented file deduplication and data-protection platform with tenant isolation, SHA-256 exact deduplication, MinHash/LSH and dHash near-duplicate detection, DLP, AES-256-GCM encryption, RBAC, and audit logging.

  • Updated Sep 13, 2026
  • JavaScript

A corpus-hygiene utility for RAG data pipelines that identifies duplicate content risk, quantifies duplication with actionable statistics, and supports controlled remediation before indexing. It enables staged audit-then-cull workflows that improve retrieval quality, reduce embedding/indexing cost, and strengthen governance in knowledge curation.

  • Updated May 5, 2026
  • Shell

Add this topic to your repo

To associate your repository with the file-deduplication topic, visit your repo's landing page and select "manage topics."

Learn more