- What happens when most published information is written with AI?
- Does AI make the web less useful as a source of knowledge?
Model collapse is what happens when a generative model is trained, generation after generation, on data that earlier models produced rather than on material people made. The rare cases and unusual phrasing at the edges of the original data disappear first, and each successive model converges on a narrower, more generic approximation of an average. It is a finding about training pipelines, named by a team of computer scientists, and it has nothing to do with how any one person uses AI.
The answer, in one line
Model collapse is the degradation that appears when a generative model is trained, generation after generation, on data produced by earlier models rather than by people.
Definition#
Model collapse: the degradation that appears when a generative model is trained, directly or indirectly, on data produced by earlier models rather than by people. The tails of the original distribution, meaning its rare and unusual examples, disappear across generations, and outputs narrow towards an increasingly generic average.
Where the tails go first#
Ilia Shumailov and five colleagues at Cambridge, Oxford, Toronto and Imperial College named the effect in a paper published in Nature in July 2024, building on a preprint they had circulated the previous year. They trained language models, variational autoencoders and Gaussian mixture models recursively, feeding each generation's output back in as the next generation's training data, and found the same pattern in all three kinds of model. The unusual examples, the ones furthest from the centre of the distribution, are the first to go. What survives is whatever was already common, so a model trained on its own descendants' output drifts towards repeating a narrower and narrower slice of what it once knew. The authors call the resulting defects irreversible once several generations have passed.
The mechanism is closer to genetic drift in a small, closed population than to a machine simply forgetting. Nothing has to go wrong for it to happen; it follows from training on a sample of a sample, with no new source of variation coming in. An Author Correction issued in March 2025 fixed a single mislabelled variable in the paper's theoretical section and changed none of its findings.
Not the same thing as knowledge collapse#
The two terms are routinely swapped, and they describe opposite subjects. Model collapse happens inside a system: a model degrades because it is trained on its own output. Knowledge collapse, a later and separate argument from Andrew Peterson, happens to the humans outside the system: a population's shared stock of knowledge narrows as cheap, AI-mediated answers pull people towards the centre of what a model already knows, at the expense of the wider and stranger material a search or a library used to surface. A model cannot choose its training data; a person choosing the easiest answer can. That difference in agency is the whole reason the two findings are not interchangeable, and a page or an argument that uses them as synonyms is describing something neither paper claims.
What nobody has measured yet#
The reason model collapse gets raised outside machine-learning circles is the worry it implies about the open web. If AI-written text keeps accumulating online, and future models are trained on web-scraped data that increasingly includes that text, the same recursive-training conditions Shumailov and colleagues built in a laboratory could start to apply to the internet at large, at whatever pace that happens. Their own paper is explicit that this is the stake: genuine, human-generated data becomes more valuable, not less, as model-generated content spreads across the pool that future models learn from.
What the paper does not do, and what nothing in this research has found anyone else doing either, is measure how far that process has actually progressed inside any deployed, real-world model, or what fraction of today's web pages are AI-written to begin with. The experiments are controlled recursive-training runs built to isolate the effect, not an audit of the live internet. Treat the mechanism as demonstrated and the scale, in the world as it stands today, as unmeasured. Anyone citing a specific percentage of the web as "already AI-written" is citing a number this paper does not contain.
Key sources
- Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R. and Gal, Y. (2024). AI models collapse when trained on recursively generated data. Nature, 631, 755-759.
- Shumailov, I., Shumaylov, Z., Zhao, Y., Gal, Y., Papernot, N. and Anderson, R. (2023). The Curse of Recursion: Training on Generated Data Makes Models Forget. arXiv:2305.17493, where the term Model Collapse first appears, recorded in the graded entry above.
Related SuperSkills research#
On the human-population version of this argument, and why it is not the same claim, knowledge collapse. On what happens when people, rather than models, converge on the same output, whether AI makes everyone think alike. On the individual-level parallel to a system losing its own edges, the Google effect.
About this research#
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. Model collapse is Shumailov and colleagues' term and finding, not his, and is defined here because the research's arguments about AI-generated content and the open web depend on using it correctly rather than as a synonym for knowledge collapse. The graded entry is in the evidence base.
Explainer · SS-2026-362 · Graded against the published rubric
Hirji, R. (2026). Model collapse. The SuperSkills evidence base, SS-2026-362. https://thesuperskills.com/research/what-is-model-collapse. Last reviewed 27 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work