Trust is earned, not given

A different perspective

2026-05-21 · Projects

AI Frontiers, part 21: Synthetic data and the model-collapse debate

Part 21from the AI Frontiers series · 65 parts in all

There is a specific kind of bad argument that technical fields produce when a real problem meets a catchy name. "Model collapse" is one of them. The underlying concern is legitimate and sharp — if each generation of models is trained on the output of the previous generation, then the training distribution drifts away from the real world, and error accumulates — but the phrase has been used to make claims ranging from trivially true to flatly wrong, and the distinction between them determines how you build a data pipeline.

The reason it is worth a whole entry is that synthetic data is not optional anymore. It is how small models get strong (part 9), how reasoning capability propagates downward (part 16), how preference data gets generated at scale, and increasingly how the last bits of real-world text get squeezed for signal. The field's ability to keep improving depends on getting this right, and the answer that has emerged is more specific than either side of the debate.

What the recursion argument actually shows

The foundational paper trained generative models on their own output, generation after generation, and observed the tails of the distribution disappearing — first the rare events, then progressively more of the distribution, producing outputs that converged toward a bland average while losing the ability to represent anything unusual (Shumailov et al., "The Curse of Recursion"). The intuition is straightforward. Sampling is not the same as the distribution. A model that samples is a lossy filter, and each pass through the filter drops the parts of reality that were least represented in the training data. Repeated, the process is a random walk away from the truth, and the popular follow-up work showed the same effect in variational autoencoders and diffusion models (Alemohammad et al.).

The Nature paper (Shumailov et al., "AI models collapse when trained on recursively generated data") is the version that travelled, and it is worth noting what it does and does not establish. It establishes that collapse happens under a specific and quite realistic regime: pure recursion, no fresh real data, no filtering, no verification. It does not establish that any model trained with synthetic data today is collapsing, and the differences between those two statements are the whole engineering question.

The counterargument, which is stronger

The definitive-sounding response came from a paper with the wonderful property of being mostly about accounting. Gerstgrasser et al. showed that the collapse result depends on replacing the training set each generation. If instead you accumulate — keep all the real data and add synthetic data to it — the model does not degrade, and their analysis suggests it can even improve. The real data anchors the distribution; synthetic data adds volume and coverage without being able to drift the target.

That result reframes everything. Model collapse is not a law of nature about synthetic data. It is a property of a specific, avoidable data pipeline: the one where someone throws away the real corpus because the synthetic one is newer or cleaner or larger. And that is exactly the mistake a well-funded team could plausibly make, because synthetic data is cheap and real data requires lawyers.

The second refinement is verification. If synthetic examples are filtered by a rule that checks whether they are actually correct — does the code run, does the proof check, does the generated question have a defensible answer — then the noise introduced by generation is removed rather than accumulated. A paper on exactly this point found that scaling up synthetic data works when verification is in the loop (Feng et al.). Which connects directly back to part 16: the reason reasoning models can be trained this way at all is that verifiable domains give you a way to tell a good generation from a bad one at scale.

The evidence that synthetic data works, early and cheap

The empirical case was made before the theory caught up. Self-Instruct (Wang et al.) bootstrapped an instruction-following dataset by having a model generate instructions and inputs from a small hand-written seed set, then filtering for validity and diversity. That recipe produced Alpaca and, more importantly, became the default way anyone with a modest budget fine-tunes an instruction follower. Magpie (Xu et al.) went further and generated instruction data from an aligned model with almost no prompt engineering, by eliciting instructions from the model's own likely continuations.

The Phi line is the cleanest demonstration that the ceiling on quality is data quality rather than data quantity. Phi-1 was trained on a deliberately small corpus of synthetic "textbook-quality" material and matched far larger models on code generation for its size (Gunasekar et al.); Phi-1.5 and Phi-2 scaled the recipe (Li et al.); Phi-3 took it to 3.8B with a long-context variant (Abdin et al.). In every case the interesting variable was not how many tokens went in but how much signal was in each token.

There is a real cost to that approach, and the Phi literature is honest about it: models trained on textbook-style synthetic material show a characteristic weakness on benchmarks measuring factual recall about the world. Density of reasoning does not buy breadth of knowledge. A pipeline optimized on synthetic data will produce a model that reasons well about what it was taught and knows less about everything else, which is a trade you should choose deliberately rather than discover.

What the pipeline actually looks like

"Synthetic data" describes a family of pipelines rather than one technique, and the differences between them explain most of the disagreement about whether the approach works.

A typical generation pipeline has five stages. Start with a seed set — a few hundred hand-written examples, or a specification, or nothing at all if the model is aligned enough to be prompted for instructions directly. Generate candidates at scale, often with diversity encouraged by sampling temperature or by conditioning on a topic or difficulty. Filter aggressively: deduplicate, reject malformed outputs, reject examples whose answers do not check out. Score what remains with a reward model or a judge. Train on the survivors, and measure the result against a real held-out set.

Every stage is where a pipeline can go wrong in a way that is invisible in the loss curve. A seed set that is not diverse produces a narrow distribution however much you generate. A filter that only checks format lets confidently wrong examples through. A judge that prefers length produces a model that prefers length. Deduplication is the most underestimated stage: near duplicates in synthetic corpora are extremely common, and a duplicated example is effectively a higher learning rate on that example, which quietly skews the distribution toward whatever the generator was most confident about. The general-purpose deduplication result from the pretraining literature (Lee et al.) applies with extra force here.

Verifying what cannot be verified

The clean case is a domain with an oracle. Code either passes the tests or does not; a mathematical answer is either right or wrong; a generated question has a defensible answer that a rule can compare against. Reasoning models live entirely in this case, which is why they work.

Most domains are messier. When there is no oracle, practitioners reach for weaker proxies, and the hierarchy of reliability is worth stating. A process reward model that scores each reasoning step is better than one that scores only the final answer, for the reasons part 11 discussed. An ensemble of judges with different prompts and different models is better than a single judge, because agreement is weak evidence of correctness. A judge given the source document and asked to check support is better than one asked whether an answer "seems good." And human spot-checks on a random sample remain the only way to detect systematic bias in an automated judge, which is the failure mode that matters because it is perfectly consistent.

The uncomfortable generalization is that synthetic data is only as trustworthy as the verification you can attach to it, and that verification quality, not generation volume, is where the returns are. A modest pipeline with an executable checker beats a large pipeline with a language-model judge scoring its own output, by a margin that surprises people who think of generation as the expensive step. It is not. Generation is cheap now. Judgment is not.

The distillation economy, and the awkward questions it raises

Distillation is synthetic data with an economic purpose: use a strong model's outputs to train a weaker one so the capability can be served cheaply. It is now the standard mechanism by which reasoning models propagate downward, and the results in part 16 show how well it works — a small model fine-tuned on a large model's reasoning traces can outperform much larger models that were never trained that way.

Three unresolved issues follow it around. The first is terms of service: whether generating training data from a commercial model's API and using it to train a competitor is permitted is a contractual question that different providers answer differently, and the answer has changed over time. The second is attribution and licensing: the traces embed the teacher's style and distribution, and there is no accepted convention for what that obliges you to disclose. The third is more technical and more interesting — distillation is a lossy channel, and the student inherits the teacher's blind spots without necessarily inheriting its calibration. A distilled model that is confidently wrong in the same way as its teacher will fool an evaluation built around the teacher.

The contamination problem underneath

There is a data-quality issue that synthetic data made worse, and it deserves more attention than the collapse debate usually gets: benchmark contamination. If a model is trained on data that was filtered, ranked, or generated with reference to a public benchmark, its score on that benchmark becomes uninterpretable. This was already a problem with scraped web data containing test sets, and generating synthetic data from a model that has seen those benchmarks compounds it. Contamination studies (Sainz et al.) documented how pervasive the effect is across standard evaluation sets, and the practical consequence is that a suspiciously large jump on a popular benchmark deserves a contamination check before it deserves a press release.

The response is not to abandon benchmarks but to treat them as probes rather than scores: hold private evaluation sets, refresh them regularly, test on data created after a model's training cutoff, and prefer tasks where you can verify correctness rather than recall. This is expensive and it is the only thing that makes a leaderboard position meaningful.

What I would actually do

Four rules, drawn from the literature and from watching pipelines fail.

Never replace real data with synthetic data. Accumulate. This one distinction decides whether the collapse result applies to you, and it is a decision made once in a data pipeline rather than tuned later.

Put a verifier in the loop. If a synthetic example cannot be checked — by a test suite, a unit conversion, a rule, a second independent model with a different prompt — it is noise with an appealing distribution. Verification is what converts synthetic data from a degradation risk into an amplifier.

Keep the mix explicit and measurable. Track the ratio of real to synthetic tokens across training, and evaluate on a held-out set that contains no synthetic data at all. The failure mode is silent, so the instrumentation has to be deliberate.

Expect the trade and choose it. Synthetic-heavy training buys reasoning density and costs world knowledge, style diversity and calibration. If your application needs a model that knows the world, weight accordingly; if it needs a model that solves a narrow class of problems well, the density is worth much more than the breadth.

My honest read of where this settles is that synthetic data is the reason the last two years looked the way they did, and the collapse literature is the reason the field will not accidentally destroy its own foundation — as long as people read it as a warning about pipelines rather than a prophecy about a technology. Every sufficiently sharp warning gets absorbed into engineering practice, and the ones that do not are the ones that were never specific. This one is specific, and the fix is boring, which is the best possible outcome for a scary-sounding problem.

Works Cited

Abdin, Marah, et al. "Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone." arXiv, 2024, arxiv.org/abs/2404.14219. Accessed 21 May 2026.

Alemohammad, Sina, et al. "Self-Consuming Generative Models Go MAD." arXiv, 2023, arxiv.org/abs/2307.01850. Accessed 21 May 2026.

Feng, Yunzhen, et al. "Beyond Model Collapse: Scaling Up with Synthesized Data Requires Verification." arXiv, 2024, arxiv.org/abs/2406.07515. Accessed 21 May 2026.

Gerstgrasser, Matthias, et al. "Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data." arXiv, 2024, arxiv.org/abs/2404.01413. Accessed 21 May 2026.

Gunasekar, Suriya, et al. "Textbooks Are All You Need." arXiv, 2023, arxiv.org/abs/2306.11644. Accessed 21 May 2026.

Lee, Kenton, et al. "Deduplicating Training Data Makes Language Models Better." arXiv, 2021, arxiv.org/abs/2107.06499. Accessed 21 May 2026.

Li, Yuanzhi, et al. "Textbooks Are All You Need II: Phi-1.5 Technical Report." arXiv, 2023, arxiv.org/abs/2309.05463. Accessed 21 May 2026.

Penedo, Guilherme, et al. "The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale." arXiv, 2024, arxiv.org/abs/2406.17557. Accessed 21 May 2026.

Sainz, Oscar, et al. "NLP Evaluation in Trouble: On the Need to Measure LLM Data Contamination for Each Benchmark." arXiv, 2023, arxiv.org/abs/2310.18018. Accessed 21 May 2026.

Shumailov, Ilia, et al. "AI Models Collapse When Trained on Recursively Generated Data." Nature, vol. 631, 2024, pp. 755–59. Accessed 21 May 2026.

---. "The Curse of Recursion: Training on Generated Data Makes Models Forget." arXiv, 2023, arxiv.org/abs/2305.17493. Accessed 21 May 2026.

Wang, Yizhong, et al. "Self-Instruct: Aligning Language Models with Self-Generated Instructions." arXiv, 2022, arxiv.org/abs/2212.10560. Accessed 21 May 2026.

Xu, Zhangchen, et al. "Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing." arXiv, 2024, arxiv.org/abs/2406.08464. Accessed 21 May 2026.