Synthetic data and model collapse
Training on model output is now standard practice, and whether it poisons the well turns on one detail people skip.
Part of the How models are made track on lAItest.
Models trained on model output collapse. Whether that happens to you depends on whether you replace the old data or add to it.
Synthetic data is not a fringe technique.
Instruction sets, preference pairs, reasoning traces, hard examples for a narrow task: a large share of post-training data is now generated by models rather than written by people. It is faster, it covers cases too rare to collect, and it can be filtered by a verifier before use. Distillation is one instance of it. So is a great deal of modern instruction tuning.
A common misconception
Commonly believed: Every generation trains on the last generation output, so quality inevitably degrades and the web eats itself.
Actually: The collapse result, Shumailov and colleagues, arXiv 2305.17493, published in Nature in 2024, is real and was measured in a specific regime: each generation trained on its predecessor output instead of the original data. Follow-up work found that when synthetic data accumulates alongside the real data rather than replacing it, the degradation is largely avoided. Replacement is the dangerous operation, not generation.
Why replacement is the part that hurts.
A model samples from what it learned, so its output under-represents the tails: the rare phrasings, the unusual cases, the thin edges of the distribution. Train the next model only on that output and the tails shrink again. Repeat, and the distribution narrows toward its own centre. Keeping the original data in the mix keeps the tails present, which is why the accumulate case behaves so differently from the replace case.
Which setup matches the regime where model collapse was demonstrated?
Answer: Each generation trains on its predecessor output in place of the original data. The result holds when synthetic data replaces real data, generation after generation. Later work found that accumulation, keeping the original corpus and adding to it, largely avoids the effect. This remains an active disagreement in the literature, so treat any flat claim in either direction with suspicion.
In one sentence
Generating training data with a model is ordinary practice. Throwing away the real data you generated it from is the part that bites.