← Research
Research · Definition

Model collapse

A model trained on its own kind's output loses the unusual cases first and narrows towards an average of itself. It is a finding about training pipelines, not about how any one person uses AI.

Last reviewed: 27 September 2026 · Next review due: 27 September 2027

Model collapse is Shumailov and colleagues' term, not this research's. This page defines it precisely, separates it from the routinely confused knowledge collapse, and names what nobody has yet measured about the open web.

Questions this page partly answersAll 1024 questions this research covers

Model collapse is what happens when a generative model is trained, generation after generation, on data that earlier models produced rather than on material people made. The rare cases and unusual phrasing at the edges of the original data disappear first, and each successive model converges on a narrower, more generic approximation of an average. It is a finding about training pipelines, named by a team of computer scientists, and it has nothing to do with how any one person uses AI.

The answer, in one line

Model collapse is the degradation that appears when a generative model is trained, generation after generation, on data produced by earlier models rather than by people.

Share as a card

Definition#

Model collapse: the degradation that appears when a generative model is trained, directly or indirectly, on data produced by earlier models rather than by people. The tails of the original distribution, meaning its rare and unusual examples, disappear across generations, and outputs narrow towards an increasingly generic average.

Share this definition as a card

Where the tails go first#

Ilia Shumailov and five colleagues at Cambridge, Oxford, Toronto and Imperial College named the effect in a paper published in Nature in July 2024, building on a preprint they had circulated the previous year. They trained language models, variational autoencoders and Gaussian mixture models recursively, feeding each generation's output back in as the next generation's training data, and found the same pattern in all three kinds of model. The unusual examples, the ones furthest from the centre of the distribution, are the first to go. What survives is whatever was already common, so a model trained on its own descendants' output drifts towards repeating a narrower and narrower slice of what it once knew. The authors call the resulting defects irreversible once several generations have passed.

The mechanism is closer to genetic drift in a small, closed population than to a machine simply forgetting. Nothing has to go wrong for it to happen; it follows from training on a sample of a sample, with no new source of variation coming in. An Author Correction issued in March 2025 fixed a single mislabelled variable in the paper's theoretical section and changed none of its findings.

Not the same thing as knowledge collapse#

The two terms are routinely swapped, and they describe opposite subjects. Model collapse happens inside a system: a model degrades because it is trained on its own output. Knowledge collapse, a later and separate argument from Andrew Peterson, happens to the humans outside the system: a population's shared stock of knowledge narrows as cheap, AI-mediated answers pull people towards the centre of what a model already knows, at the expense of the wider and stranger material a search or a library used to surface. A model cannot choose its training data; a person choosing the easiest answer can. That difference in agency is the whole reason the two findings are not interchangeable, and a page or an argument that uses them as synonyms is describing something neither paper claims.

What nobody has measured yet#

The reason model collapse gets raised outside machine-learning circles is the worry it implies about the open web. If AI-written text keeps accumulating online, and future models are trained on web-scraped data that increasingly includes that text, the same recursive-training conditions Shumailov and colleagues built in a laboratory could start to apply to the internet at large, at whatever pace that happens. Their own paper is explicit that this is the stake: genuine, human-generated data becomes more valuable, not less, as model-generated content spreads across the pool that future models learn from.

What the paper does not do, and what nothing in this research has found anyone else doing either, is measure how far that process has actually progressed inside any deployed, real-world model, or what fraction of today's web pages are AI-written to begin with. The experiments are controlled recursive-training runs built to isolate the effect, not an audit of the live internet. Treat the mechanism as demonstrated and the scale, in the world as it stands today, as unmeasured. Anyone citing a specific percentage of the web as "already AI-written" is citing a number this paper does not contain.

Key sources

On the human-population version of this argument, and why it is not the same claim, knowledge collapse. On what happens when people, rather than models, converge on the same output, whether AI makes everyone think alike. On the individual-level parallel to a system losing its own edges, the Google effect.

About this research#

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. Model collapse is Shumailov and colleagues' term and finding, not his, and is defined here because the research's arguments about AI-generated content and the open web depend on using it correctly rather than as a synonym for knowledge collapse. The graded entry is in the evidence base.

How this research works  ·  Reviewed quarterly  ·  Found an error? Tell me and it is corrected on the page.

Explainer · SS-2026-362 · Graded against the published rubric

Cite this page

Hirji, R. (2026). Model collapse. The SuperSkills evidence base, SS-2026-362. https://thesuperskills.com/research/what-is-model-collapse. Last reviewed 27 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

What is model collapse?

Model collapse is the degradation that appears when a generative model is trained, generation after generation, on data produced by earlier models rather than by people. Shumailov and five colleagues named it in a paper published in Nature in July 2024, after training language models, variational autoencoders and Gaussian mixture models recursively on their own descendants' output. In every case the rare and unusual examples in the data disappeared first, and each generation converged on a narrower, more generic approximation of the original.

Is model collapse the same as knowledge collapse?

No, and the two are routinely swapped for each other. Model collapse happens inside a system: a model degrades because it is trained on its own kind's output, and it cannot choose otherwise. Knowledge collapse, a separate and later argument from Andrew Peterson, happens to the humans outside the system, as a population's shared stock of knowledge narrows because cheap AI-mediated answers pull people towards the centre of what a model already knows. A model has no choice in its training data; a person choosing the easiest answer does, which is the distinction that keeps the two findings from being interchangeable.

Does model collapse mean the internet is already getting worse?

The paper does not measure that, and nothing found for this research does either. Shumailov and colleagues ran controlled recursive-training experiments to isolate the mechanism; they did not audit how much of the live web is AI-written or how far any deployed model has actually drifted. Their own point is that human-generated data becomes more valuable as AI-generated content spreads through what future models are trained on, which is a warning about a direction rather than a measured amount. A specific percentage of the web described as 'already AI-written' is not a figure this paper contains.

Where did the term come from?

Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot and Ross Anderson introduced Model Collapse in a preprint submitted in May 2023, The Curse of Recursion: Training on Generated Data Makes Models Forget, and then published the peer-reviewed version in Nature in July 2024 under the title AI models collapse when trained on recursively generated data. An Author Correction in March 2025 fixed a single notation error in the paper's theoretical section and changed no finding.

In this hub

Definitions

The terms this field uses, defined against their primary sources.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory  ·  Boards  ·  Enquire