When AI Eats Itself: Understanding Model Collapse and AI Cannibalism

There is a snake in the AI industry, and it is eating its own tail.
For years, the recipe for building better AI seemed simple: scrape more data from the internet, train a bigger model, repeat. But we have reached a strange turning point. A huge portion of the internet is now written by AI. And the next generation of AI is being trained on that very content. The models are eating themselves.
This is not a metaphor. It has a name, and it is a formally proven problem called model collapse.
What the Researchers Found
In July 2024, a team led by Ilia Shumailov at Oxford University, working with researchers from Cambridge and other institutions, published a landmark paper in the journal Nature. Their finding was stark and simple: if you train AI models on data generated by other AI models, they will break down in an irreversible way.
They called this effect "model collapse".
The study showed that this does not just happen to one type of model. They observed the collapse in large language models (LLMs), in image models called variational autoencoders, and even in simple statistical models called Gaussian mixture models. The problem is baked into the very nature of how these systems learn.
Why Does Collapse Happen?
To understand why, you need to think about what a model actually learns. It is not learning "facts." It is learning statistical patterns from enormous piles of data. It learns what words tend to follow other words, what ideas tend to appear together.
A healthy dataset is full of variety. It contains common ideas that appear all the time. But it also contains rare things: minority languages, unusual turns of phrase, niche knowledge, strange edge cases. Researchers call these rare parts the "tails" of the data distribution.
Here is the problem: every model slightly overproduces the common stuff and underproduces the rare stuff. It is a small bias, but it exists in every generation of the model.
Now imagine you take that output — which is slightly skewed toward the average — and you train the next model on it. The rare content shrinks even more. You do this again, and again. The tails of the distribution disappear entirely.
What you are left with is a model that has forgotten the richness of the real world. It only knows the middle, the average, the bland. Its outputs become repetitive and eventually nonsensical.
The Jackrabbit Problem
The researchers gave a vivid demonstration of this. They took a small model called OPT-125m and fed it a passage of text about 14th-century church towers.
Then they let it generate new text. They trained a new model on that output. Then another. And another. They did this for nine generations.
By generation nine, the model had completely forgotten about architecture. When asked about medieval church towers, it started listing species of jackrabbits — including fictional ones like "blue-tailed jackrabbits".
The model did not just get a fact wrong. It had lost the entire conceptual grounding of the original topic. It was hallucinating entirely.
This Is Happening Right Now
This is not a theoretical problem for the future. It is a problem for right now.
A study from the company Graphite analyzed over 55,000 English-language articles published between 2020 and 2026. They found that in the first quarter of 2026, an estimated 49.9% of sampled articles were classified as mostly AI-generated.
In late 2024, that figure was around 48%. By late 2025, AI-generated articles briefly surpassed human-written ones.
Other studies paint a similar picture. One analysis of over 1.2 million webpages from late 2025 to early 2026 found that roughly 35% contained wholly or partially AI-generated content. Estimates suggest that 30–40% of the active web corpus is now synthetic.
This matters because those same models are trained on web-scraped data. They are increasingly ingesting content that other models wrote. The contamination is not a future scenario. It is the current state of the training pipeline.
Can It Be Fixed?
The research is not all doom. But the fixes are not simple.
The most basic idea is to mix real data with synthetic data. A paper presented at an IEEE conference in late 2025 proposed a simple strategy: combine synthetic data and human-generated data so that the model does not degenerate. They derived mathematical conditions on the ratio of synthetic to human data needed to keep the system stable.
But there is a catch. The same paper notes that on a small GPT model, enriching synthetic data with just a small amount of human data may not be enough to prevent collapse. The balance matters.
Another approach is to use detectors. Researchers at University College London trained a machine-generated text detector and proposed a method to up-sample likely human content in the training data. They found that this not only prevented model collapse but actually improved performance compared to training on purely human data.
A third approach is to monitor entropy. As a model collapses, the entropy of its training data declines. The model stops generating novel content and starts memorizing its training samples. Researchers have proposed using this entropy decline as an indicator of degradation and selecting data to preserve diversity.
A fourth idea is verification. Some researchers have shown that injecting information through an external verifier — whether a human or a better model — can prevent synthetic retraining from causing collapse.
The common thread in all these solutions is the same: human-generated data is irreplaceable. The original paper put it plainly: "the value of data collected about genuine human interactions with systems will be increasingly valuable in the presence of LLM-generated content in data crawled from the Internet".
What This Means for You
If you are building or training models, the lesson is straightforward: know where your data comes from. If you are scraping the web post-2022, you are training on synthetic content. Monitor your model's diversity and entropy during training.
If you are working with data pipelines, understand that genuine human data is becoming a strategic asset. The value of real conversations, original writing, and real-world observations is rising precisely because it is becoming scarce.
And if you are watching the industry broadly, the "free lunch" of scraping unlimited web data is ending. Sustainable AI development will require deliberate investment in data provenance, curation, and verification.
The snake can stop eating its tail. But it has to choose to do so before the tails disappear entirely.
Sources
Shumailov, I. et al. "AI models collapse when trained on recursively generated data." Nature 631, 755–759 (2024).
Graphite study on AI-generated articles, reported by Search Engine Land (2026).
Data Science Dojo, "AI Cannibalism: How Models Are Eating Themselves Into Collapse."
IEEE conference paper, "Preventing Model Collapse when Training LLMs with Synthetic Data" (2025).
UCL research on machine-generated text detection (2025).
NeurIPS paper, "A Closer Look at Model Collapse" (2025).
Cover image source: Data Science Dojo



