Est.

Deduplication and Near-Duplicate Detection in Document Corpora

Removing duplicates cuts training tokens by 13-41% while preventing memorization.

Columnist · · 10 min read
Cover illustration for “Deduplication and Near-Duplicate Detection in Document Corpora”
RAG ingestion and chunking strategy · September 30, 2026 · 10 min read · 2,205 words

Half or more of a raw corpus can be duplicate or near-duplicate documents. Deduplication is a foundational step, not cleanup work bolted onto a pipeline. That foundational step determines what a model really learns The RefinedWeb Dataset for Falcon LLM. People using C4 spotted a single 61-word English line copied over 60,000 times. This is a widespread, systemic pattern. That’s the result when scraped web material gets folded into training data with no check for redundancy.

Skipping this step causes problems in multiple ways. If deduplication is skipped, more than 1% of a model's unprompted output turns into text copied verbatim from training data, its most obvious effect. Subtle and arguably more damaging, this form of contamination undermines the yardstick itself. Over 4% of a typical validation collection may involve train-test overlap. Benchmark results can look higher than genuine generalization supports. Also, when training data holds duplicate sequences, membership inference and extraction work more often, so a data hygiene concern becomes a privacy one.

The gain from dealing with this is just as clear. With solid deduplication, networks hit the same or higher accuracy while cutting training tokens by 13 to 41%, and they rarely output memorized material. Those cuts to compute are real savings. Duplicates range from exact matches to near-copies and differently worded content that means the same thing semantically, so no one approach catches them all. Picking a sequence of methods is a real engineering call with tradeoffs, which is what the rest of this article covers.

The fundamental split between exact and near-duplicate detection

Exact deduplication solves only this: when two documents match byte-for-byte, retain one copy and delete the others. Finding near-duplicate content solves a messier challenge. Documents have substantial content in common but vary in forms hashing misses, like OCR noise, different formatting, edits, cross-format republication, or paraphrasing.

This difference between the two matters in practice. On WildChat prompts in a real-world check, byte-exact deduplication found duplicates at only 5.81%, while MinHash-LSH at the 0.9 Jaccard threshold found 31.32% Zyda dataset. Switching methods on the same data lifts coverage past fivefold. The two approaches each answer different questions. Each approach is answering different questions, so pipelines doing real deduplication run both in sequence.

Even before then, normalization handles a surprising share. Stray whitespace, encoding quirks, and letter casing mean hashing raw documents will skip duplicates a person would spot right away. Fixing the normalization step changes how many duplicates surface, and with limited corpora it can count for more than the algorithm picked later. Substring-level deduplication finds repeated text even when the rest of the file changes. Lee et al. focused their pipeline on pulling out every duplicate substring above 50 BPE tokens, working beneath the document level so repeated passages get flagged regardless of what surrounds them.

Diagram: Same Data, Fivefold More Duplicates Found. Visualizes: Show a magnitude contrast between two deduplication methods applied to the same WildChat prompt corpus: byte-exact deduplication found duplicates in 5.81% of documents, while…

Exact deduplication in practice, from hash sets to suffix arrays

At limited scale, exact deduplication ranks among the simplest things to build. Normalize the text, run MD5 or SHA-256 on it, keep the hash in a lookup, and discard what collides GoPenAI. Lookup is O(1), and at this scale of a few thousand documents, the memory cost hardly registers GoPenAI.

When a corpus no longer fits in memory, the workflow moves to streaming with a disk-based hash such as LevelDB, LMDB, or Redis. With document counts in the millions, fingerprints are still manageable though the source material isn't. At still bigger scale, Bloom filters provide a probabilistic shortcut: Dolma applies them for document-level and paragraph-level deduplication, allowing a modest false-positive share for a memory footprint dramatically below an exact hash table.

Substring-level deduplication calls for a different data structure, since checking every document pair for shared substrings runs in quadratic time and breaks down at corpus scale. To do this, concatenate all the corpus as one long sequence, create a linear suffix array for it, then scan that array once for duplicate substrings. RefinedWeb and MADLAD-400 run on this approach SlimPajama deduplication study Falcon RefinedWeb. After MinHash, RefinedWeb still ran exact substring removal and shrank an already-deduplicated corpus by nearly another 40% SlimPajama deduplication study Falcon RefinedWeb. Exact methods are quick, reliable, and simple to check. Use them first in any pipeline, but alone they miss plenty of almost-duplicate text.

Shingling, MinHash, and LSH: how fuzzy near-duplicate detection scales

Shingling is the base that both MinHash and SimHash build on. A shingle means a contiguous sequence of tokens, and shingling turns each document into overlapping n-grams so both algorithms can compare real data. Checking exact Jaccard similarity for every pair of documents in a large corpus is quadratic, and at scale, brute checks become a non-starter because of the combinatorial blowup.

MinHash avoids the exact math, estimating Jaccard similarity as a shortcut. Hash each document's n-grams with multiple hash functions to produce a compact code. Those signatures are put into bands, then hashed to buckets; any two documents in one bucket form a pair for closer review. Comparing only those pairs turns a quadratic mess into work that stays tractable on real corpus scale. Surviving candidate pairs then merge into duplicate clusters via a disjoint-set structure, known as union-find.

These parameter settings aren't just cosmetic. They change what "duplicate" even means for a given corpus. Penedo's team applied 5-grams across its 2023 and 2024 studies; Gao's 2020 team and Soldaini's in 2024 went with 13-grams. Using bigger n-grams means two documents need more stretches of exact overlap to qualify as matches, which shapes recall. Even more important is the Jaccard threshold Zyda dataset. SlimPajama started with 605 billion raw tokens, but the 80% threshold left only 355 billion in the corpus, and a 40% threshold dropped it to 242 billion: picking that threshold shifted things by over 100 billion tokens SlimPajama deduplication study Falcon RefinedWeb Institutional Books 1.0. At 40%, the filter mostly flags boilerplate and content that's been reformatted, keeping the sense but not the exact wording SlimPajama deduplication study Falcon RefinedWeb. At 80%, what's flagged is basically the same at the level of individual characters. The threshold controls how much gets caught semantically.

This pattern shows up in real datasets. Domain matters too: corpus HC4 used 256 hashes for each healthcare document, 5-grams and threshold of 0.85, tuned to the noise profile of medical and biomedical writing GoPenAI. For cross-corpus deduplication, the 1.3T token dataset Zyda applied MinHash to its corpora at 80% and 40% thresholds, dropping the Pile-Uncopyrighted corpus from 253 billion raw tokens down to 83 billion using that 40% threshold GoPenAI.

Diagram: How the Jaccard Threshold Shapes What Survives. Visualizes: Show how a single threshold setting dramatically changes corpus size starting from 605 billion raw tokens: an 80% Jaccard threshold left 355 billion tokens; a 40% threshold left…

SimHash versus MinHash preference

SimHash uses a different geometric method. It maps like documents into nearby hashes by Hamming distance, defining close in a different way from estimating Jaccard similarity.

The bigger surprise was where it broke. SimHash behaves more or less as bag-of-words does, and long documents get flagged as alike even without meaningful content in common. ROOTS ran into this problem: false positives clustered in long documents, leading them to keep every document once its near-duplicate cluster exceeded 6,000 characters BigScience ROOTS project.

There's a clear real-world engineering point here. SimHash runs fast with little memory, but sensitivity to document length calls for a safeguard bolted on when the corpus has long-form material. It usually shines with shorter, similar documents, where size alone causes the flagging. After exact deduplication, the 1.6 TB corpus spanning 59 languages, BigScience ROOTS, applied SimHash on OSCAR with 6-grams at a Hamming distance threshold of 4 BigScience ROOTS project. They found near-duplicates in 0.7% of documents on the whole, though it ran from 0.07%–2.7% across languages, a small number applied to a corpus that had already been exact-deduplicated BigScience ROOTS project.

Embedding-based and soft deduplication for semantic similarity

Both MinHash and SimHash have a known weakness. They skip documents that carry the same point in different phrasing, like paraphrases or restructured sentences or translated content, every time the sense overlaps yet the surface tokens differ. Catching that means understanding semantics, not just counting shingles.

At ICLR 2023, an evaluation measured hashing, n-gram overlap, a bi-encoder trained contrastively, and a rerank-style bi-encoder-plus-cross-encoder pipeline using a 27,210-document historical corpus containing 122,876 labeled duplicate pairs. The neural approaches beat the n-gram methods by a wide margin, and the bi-encoder scaled across 10 million articles on one GPU within hours.

SemDeDup turns this concept into a pipeline. Each document or image is encoded into its own embedding, K-means groups the embeddings, and duplicate search runs per cluster, not over every embedding, keeping cost down. In each cluster, items over the cosine similarity threshold count as semantic duplicates, and the point nearest the cluster centroid survives. It applies to words and images, so it's a good choice for datasets that are multi-modal.

SoftDedup works from a different philosophy: without deleting a thing, it downweights sampling of high-commonness data, relying on an n-gram to gauge each passage's frequency. It preserves all document entries in the dataset, cuts training time 26% or more at similar perplexity, and lifts few-shot downstream accuracy 1.77% with the same training spend. Staying non-destructive matters most when deletion itself threatens rare-language corpora or specialized domains where giving up a single document carries a real cost. Silcock et al. applies neural methods using both bi-encoder and cross-encoder setups.

Choosing a method: the precision, recall, and compute tradeoff space

Here, each option balances the same axes: recall (catching near-duplicates), precision (avoiding flagging different documents as duplicates), and compute cost. Exact hashing is one end of the range: precision never slips, recall stays at zero unless files are byte-exact, and its cost is negligible. That alone makes it worth using first in nearly every pipeline.

MinHash-LSH occupies the middle ground that most production systems actually live in. Threshold and n-gram settings make its precision and recall tunable, compute remains tractable at corpus scale, and GPT-3, the Pile, RefinedWeb, SlimPajama, Zyda, HC4, plus Meltemi use it most. SimHash runs faster and takes up less memory footprint, sacrificing some tunability at the cost of that length-sensitivity covered earlier. Neural and embedding methods sit at the farthest point: top recall for true semantic duplicates, though at meaningfully more compute cost, and that cost pays off mostly where n-gram methods demonstrably fail on noisy OCR, duplication across languages, or content that has been heavily reformatted. SoftDedup is the right call specifically when deletion itself is unacceptable, whether for regulatory reasons, provenance requirements, or the simple scarcity of data in an underrepresented language.

None of the recent big corpora stick to a single method. RefinedWeb used MinHash before exact substring removal. ROOTS applied exact deduplication ahead of SimHash. Meltemi applied MinHash per dataset, then deduplicated again between datasets. The WildChat comparison above finds MinHash-LSH and exact methods flag different duplicate sets, and relying on either one leaves the rest still untouched in the corpus.

What real corpus pipelines look like end to end

MinHash ran first in the pipeline as NEARDEDUP to flag approximate near-duplicates, with exact substring deduplication, EXACTSUBSTR, applied afterward, removing nearly 40% more from the already-reduced corpus. Together, exact and fuzzy deduplication cut the starting dataset by roughly 50%. For the long-document safeguard, documents over 6,000 characters got excluded from discard decisions because of false-positive concerns. Zyda is the 1.3T token dataset presented to illustrate this.

Dolma handled exact deduplication at document and paragraph level with probabilistic Bloom filters for memory efficiency over a large corpus. Meltemi, the Greek LLM, ran two-stage MinHash: first inside every constituent dataset, then over the merged dataset after concatenation, to catch intra and inter-source duplication. HC4 tuned for the noise common in medical and biomedical text: MinHash LSH with 256 hashes per document, using 5-grams and a 0.85 threshold. OSCAR underwent exact deduplication first. SimHash ran on 6-grams with a Hamming distance cutoff of 4 to catch near-duplicates.

That's deduplication in an archival setting, where the definition of "duplicate" has to be built around the domain rather than borrowed off the shelf. In every case, the same pattern shows up: no team ships one algorithm and treats the job as finished. They run methods in sequence, with each step catching what the previous one missed. Falcon RefinedWeb is presented as a case, using about 5,000GT extracted data from CommonCrawl, with 600 billion tokens made public. BigScience ROOTS, comprising 1.6 TB spanning 59 languages, gets presented as one case.

Engineering realities: cost, tooling, and GPU acceleration

Deduplication compute costs grow with the corpus, so it pays to price them out with real numbers. On a 32-core CPU box with 256 GB of RAM, DataTrove finishes MinHash over a corpus of roughly 1–10 billion tokens (about 100 GB) in approximately 4–8 clock-hours, costing roughly $15–$30. At this scale, skipping deduplication on cost grounds is tough to defend.

At this scale, 10 to 100 billion tokens, spanning 1 to 10 terabytes in total, the tooling moves to DataTrove on SLURM cluster infrastructure, or NeMo Curator with multi-GPU support. Wall-clock spans 12 to 48 hours, with deduplication costing roughly $100. Training a bloated, duplicate-rich corpus for additional epochs would burn far more compute and amplify memorization risk than the dedup pass. Now that the tooling has matured, deciding to deduplicate is no longer the hard part. It's which mix of methods, in what staged sequence, genuinely fits the corpus's noise profile.

Sources

  1. The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only
  2. The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset
  3. Why Deduplication Is the Most Underestimated Step in LLM Pretraining And What It Costs You to Get It Wrong | by Ibrahimdaud | GoPenAI
  4. Institutional Books 1.0: A 242B token dataset from Harvard Library's collections, refined for accuracy and usability

More in RAG ingestion and chunking strategy