AI Frontiers, part 40: Data curation β deduplication, filtering, and why clean beats big
Part 40from the AI Frontiers series · 65 parts in all
Every entry in this series that discusses capability eventually runs into the same fact: the differences between models trained on similar compute are frequently explained by their data, and by the pipeline that produced it rather than by the corpus it came from. A web scrape is a raw material. What determines a model's ceiling is what survived deduplication, what was filtered out, what was echoed more than once, and what was never there at all.
This is the least glamorous part of the stack and it is the one with the highest leverage. It is also where a great many unstated value judgments get encoded, which is why the interesting literature in this period is as much about what filtering does to underrepresented content as about what it does to model quality.
The pipeline, stage by stage
A modern curation pipeline has a recognizable sequence, and each stage has a characteristic failure mode.
Extraction and boilerplate removal. Raw HTML is stripped of navigation, advertising, templates and markup. This sounds trivial and is not: the first-generation pipelines that did it badly left a great deal of repeated footer text in the corpus, which teaches a model to generate repeated footer text.
Language identification. Documents are classified by language and routed accordingly. The failure mode is systematic under-detection for languages that are less represented in the identifier's training data, which compounds the underrepresentation that already exists.
Quality filtering. Heuristics β average word length, ratio of punctuation to text, presence of common stopwords, repetition of n-grams β remove machine-generated spam, link farms and near-garbage. These heuristics were designed for English and generalize unevenly, which is the subject of a paper worth taking seriously: Gururangan et al. showed that standard quality filters systematically exclude text from minority dialects and non-dominant English varieties, because dialectal writing looks statistically abnormal to a filter built on standard English. The filter is doing exactly what it was asked to do. Nobody asked whether that was the right question.
Deduplication. The stage with the clearest measured benefit, discussed below.
Diversity and deduplication of topic. Beyond removing identical documents, a corpus can be correlated in less obvious ways β the same news story reproduced across thousands of sites, the same technical documentation in every repository. Clustering and de-duplicating at the semantic level (Tirumala et al.) measurably improves training efficiency by reducing the effective repetition rate.
Mixing. Finally, the ratios between sources β web text, books, code, scholarly articles, mathematics β are themselves a hyperparameter, and getting them wrong is a common cause of a model that is good at one thing and unexpectedly bad at another.
Deduplication, and why the effect is so large
The result that made the industry take this seriously was the finding that deduplicating a training corpus improves a model's quality at fixed compute, and that the improvement was large enough to be worth the substantial engineering cost (Lee et al.). The mechanism is straightforward once stated.
A training objective weights tokens by frequency. A document that appears ten times is effectively trained on at a rate ten times higher than a document that appears once, which means the model implicitly allocates capacity to whatever happened to be duplicated by the crawl β mirrors, syndicated content, repositories that vendored a library, boilerplate reproduced across millions of pages. The fact that the model spends capacity memorizing a privacy policy that appears on every page of a website is not a capability, it is a data defect. Deduplication removes it, which both improves generalization and reduces verbatim memorization, the phenomenon Carlini et al. documented in the privacy literature.
Two practical subtleties. Exact-match deduplication catches only identical strings, and the web is full of near-duplicates β pages with differing timestamps, ad blocks or navigation. Near-duplicate detection needs locality-sensitive hashing or MinHash-style methods, which scale acceptably but not trivially. And the unit of deduplication matters: removing duplicate documents leaves repeated passages inside distinct documents, which is why line-level and paragraph-level deduplication (as in the Gopher pipeline, Rae et al.) were used alongside document-level methods.
The second subtlety is the epoch question. If you are data-constrained, repeating epochs can help β Muennighoff et al. found that repeating a corpus up to about four times yields returns comparable to new data, with diminishing benefit after that. Curating the repeat order and repeating the highest quality material rather than the average is worth more than the general recipe suggests.
Open corpora, and what the ecosystem gained by publishing them
The open model ecosystem's most underrated artifact is not a model. It is the set of openly documented datasets and pipelines that let others reproduce the raw material.
The Pile (Gao et al.) established the practice of publishing a diverse, documented corpus with its composition visible. RefinedWeb (Penedo et al.) demonstrated that filtering web data aggressively and documenting the process could produce a corpus competitive with curated mixtures, and made the pipeline itself part of the contribution. The BigScience ROOTS corpus (LaurenΓ§on et al.) assembled multilingual data across many languages with per-language documentation. And Dolma (Soldaini et al.) went further than any of its predecessors by releasing the tooling alongside the corpus, so that the pipeline was reproducible rather than merely described.
Two consequences followed that matter beyond data. First, the gap between an organization with a good pipeline and one without narrowed, because the recipes were public; the remaining advantage lies in evaluation and iteration speed rather than in secret sauce. Second, the provenance conversation became possible at all. You cannot audit what nobody documented, and the licensing and provenance work described in part 37 depended on the existence of datasets whose composition somebody had taken the trouble to write down.
What filtering encodes, and who decides
The uncomfortable part of this area is that every filter is a policy. Deciding that text must have a certain ratio of stopwords, or a minimum average word length, or a specific syntactic profile is deciding what kind of writing deserves to be learned from β and those decisions were made, in most cases, by people optimizing an English-language benchmark.
The evidence that this has consequences is now substantial. Quality filters disproportionately remove text in African American English and other non-dominant varieties, which means a model trained on a filtered corpus has less exposure to how a large population actually writes. Multilingual coverage is systematically skewed toward languages with abundant high-resource data. Content by and about communities that are underrepresented on the web is underrepresented in the model, and it is underrepresented twice, because the filters remove some of what little there was.
This is not an argument for abandoning filtering. Unfiltered corpora produce worse models by every measure, including on the fairness evaluations that matter. It is an argument for measuring what the filter removed, documenting the choice, and treating the composition of the corpus as something to be evaluated rather than merely executed. Very few organizations can answer the question of whether their filtering disproportionately removed one variety of English over another, and the ones that can are the ones whose models behave more consistently across the populations they serve.
Building a pipeline you can defend
The engineering discipline that separates a research pipeline from an operational one is reproducibility, and in this area it is not an abstract virtue. When a model behaves unexpectedly six months later, the question is always what data it saw, and the answer requires that the corpus have been versioned, that its contents be addressable, and that the processing code still run.
Four practices make this possible. Store the raw crawl rather than only the processed output, so that a change to the filter can be applied retrospectively rather than requiring a re-crawl. Version the pipeline alongside the data, so that the specific combination that produced a checkpoint is recorded. Hash documents rather than tracking them by URL, since the same content appears at many addresses and the address is frequently the least stable part. And write a data card β composition by source, license categories, languages, known gaps, and the filter thresholds that were applied β so that the corpus can be described to someone who did not build it.
The last item is the one that pays off twice. It makes the licensing and provenance review in part 37 answerable, and it makes the fairness question above tractable, because the only way to know what the filter removed is to have recorded what it removed.
Synthetic data as a curation tool
Models do not only consume curated data; they produce it, and by 2024 the most practical use of a language model in a data pipeline was not generating training examples from scratch but transforming and auditing existing ones.
Three patterns were in wide use. Rewriting documents into a more consistent format β normalizing structure, removing conversational cruft, converting prose into question-answer pairs β improves effective quality at almost no compute cost and preserves the factual content of the original. Using a judge model to score documents against an explicit quality rubric extends the filter beyond what heuristics can express, at the cost of inheriting the judge's biases. And detecting the specific pathologies that ruin a corpus β boilerplate, machine-translated garbage, low-information content that satisfies every heuristic β is a task models are better at than regular expressions.
The caveat is the one from part 21, applied to filtering rather than training. A model-driven filter that removes the same kind of text a model-driven generator produces will systematically favor its own distribution, which narrows the corpus even as it appears to improve it. Measurement, again, is the only defense: keep the removed documents and sample them periodically to see what you threw away.
The quality-quantity trade, made explicit
Every filter has a cost, and it is worth stating as an equation. Tightening a threshold improves the average quality of the retained text and reduces the number of tokens available, and the model is trained on the product of the two. Move the threshold too far and you have a cleaner corpus that the model has seen less of; move it too little and you have a large corpus full of material that wastes capacity.
The optimum is not a fixed point and it depends on scale. A model trained once over a small corpus benefits much more from aggressive filtering than a model on its fourth epoch of a huge one, because the cost of discarding tokens is what changes. This is why the same pipeline that produces an excellent 7B model can produce an undertrained one at 70B, and why the recipe has to be re-tuned as scale increases rather than carried over.
The practical approach is to treat the threshold as a hyperparameter and measure the loss curve at small scale under several settings. It is a cheap experiment that most teams skip, and skipping it is how an organization ends up with a filtering decision inherited from a tutorial and never revisited.
Why this remains the highest-leverage work
Three reasons, in ascending order of importance.
Data work compounds while architecture does not. A better attention variant is replaced by the next one; a curation pipeline improves every model trained with it, for years. The organizations that invested early in data infrastructure are still ahead, and the gap is not closing through architecture.
Data work is where the frontier's remaining headroom sits. Compute is expensive and scaling is capital-limited. The same model trained on better data is the cheapest capability improvement available, and the Phi line's demonstration that a small model on curated material could match much larger models trained on the ordinary web (Gunasekar et al.) is the clearest evidence that data quality is not a second-order effect.
And data work is where the field's unresolved questions live: which rights the corpus implicates, which languages it serves, which communities it represents, whether synthetic material should be accumulated rather than substituted, and who is accountable when a filter quietly decides that some people's writing is not worth learning from. Those are not questions an architecture can answer, which is the strongest possible argument that the interesting work in this field is not all in the model.
Works Cited
Carlini, Nicholas, et al. "Extracting Training Data from Large Language Models." arXiv, 2020, arxiv.org/abs/2012.07805. Accessed 6 June 2024.
Dodge, Jesse, et al. "Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus." arXiv, 2021, arxiv.org/abs/2104.08758. Accessed 6 June 2024.
Gao, Leo, et al. "The Pile: An 800GB Dataset of Diverse Text for Language Modeling." arXiv, 2020, arxiv.org/abs/2101.00027. Accessed 6 June 2024.
Gunasekar, Suriya, et al. "Textbooks Are All You Need." arXiv, 2023, arxiv.org/abs/2306.11644. Accessed 6 June 2024.
Gururangan, Suchin, et al. "Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection." arXiv, 2022, arxiv.org/abs/2201.10474. Accessed 6 June 2024.
LaurenΓ§on, Hugo, et al. "The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset." arXiv, 2023, arxiv.org/abs/2303.03915. Accessed 6 June 2024.
Lee, Katherine, et al. "Deduplicating Training Data Makes Language Models Better." arXiv, 2021, arxiv.org/abs/2107.06499. Accessed 6 June 2024.
Muennighoff, Niklas, et al. "Scaling Data-Constrained Language Models." arXiv, 2023, arxiv.org/abs/2305.16264. Accessed 6 June 2024.
Penedo, Guilherme, et al. "The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only." arXiv, 2023, arxiv.org/abs/2306.01116. Accessed 6 June 2024.
Rae, Jack W., et al. "Scaling Language Models: Methods, Analysis and Insights from Training Gopher." arXiv, 2021, arxiv.org/abs/2112.11446. Accessed 6 June 2024.
Soldaini, Luca, et al. "Dolma: An Open Corpus of Three Trillion Tokens for Language Model Pretraining Research." arXiv, 2024, arxiv.org/abs/2402.00159. Accessed 6 June 2024.
Tirumala, Kushal, et al. "D4: Improving LLM Pretraining via Document De-Duplication and Diversification." arXiv, 2023, arxiv.org/abs/2308.12284. Accessed 6 June 2024.