Often enough that it now has two counts. In published, peer-reviewed medicine, one paper in 277 on PubMed cited a study that does not exist in early 2026, against one in 2,828 in 2023. In court, 2,022 decisions worldwide have recorded somebody filing material that was never written. Both counts come from people whose job is checking sources, working without a deadline at two in the morning, and both are floors rather than estimates.
Definition#
A fabricated citation: a reference produced in the correct form, with a plausible author, a real journal and a fitting year, for a work that was never published. It is distinct from a misquotation, because there is no document to have misquoted.
One paper in 277#
An audit across 2.5 million biomedical papers found 4,046 fabricated references across 2,810 papers. The rate of papers carrying at least one rose from 1 in 2,828 in 2023, to 1 in 458 in 2025, to 1 in 277 in the first seven weeks of 2026: roughly twelvefold in three years.
Review articles ran 57 per cent higher than other paper types, which makes sense and is the part that should worry a student most. A review is where somebody summarises a literature they have not personally run, which is structurally the same act as asking a model what a field says.
These are papers that went through peer review, written by researchers, in journals indexed on PubMed.
Two thousand court decisions, and most of them are not lawyers#
The AI Hallucination Cases database tracks legal decisions in which a court found that generative AI had produced hallucinated content in material put before it. Read at source on 6 September 2026, it holds 2,022 decisions.
By jurisdiction: the United States 1,379, Canada 217, Australia 110, the United Kingdom 69, Israel 57, Brazil 41, India and Italy 15 each, France 13, Germany 11, and roughly thirty further countries.
The figure most reporting leaves out is who filed the material. Pro se litigants, people representing themselves, account for 1,163 of the recorded parties. Lawyers account for 805. Judges appear 31 times, experts 15, prosecutors 5.
Read those two numbers together. The professional half of that count is people with a duty to check, insurance to lose and a regulator watching. The larger half is people with none of those things, doing what a student does: reaching for a tool under pressure, in an unfamiliar system, with no way of telling a real citation from an invented one.
Why both numbers are floors#
Neither count measures how often models fabricate. Both measure how often somebody was caught.
The court database says so in its own words: it tracks decisions and "does not track the (necessarily wider) universe of all fake citations or use of AI in court filings". A fabricated citation that nobody checks is absent by construction. The same applies to the PubMed audit, which finds what an automated check can find in a published record.
Growth in either count also mixes three things: more fabrication, better detection, and better compilation. The court database is updated daily and its author calls it a work in progress, so any figure taken from it has to carry the date it was read. This page read it on 6 September 2026 and the number will be higher by the time you do.
The property that makes this different from ordinary error#
A wrong citation has always been possible. What is new is the form the wrongness takes.
A model does not produce an obviously broken reference. It produces one in exactly the shape a real citation takes, because that shape is what it learned. Plausible author, real journal, year that fits the argument, page range in the right format. Every instinct a reader has for spotting a weak source was trained on sources that exist, and none of those instincts fires.
That is why the first move of the source rule is to establish that the source exists, before any judgement about whether it is good. Caulfield's SIFT and the CRAAP test both begin at the second question, because both were written when the first one did not need asking.
What a student should take from two numbers about doctors and lawyers#
Not that AI is untrustworthy, which is too broad to act on. Something narrower and more useful: the people in these counts were not careless, and being careful is not the defence.
A researcher writing a review and a litigant filing a brief both did the thing that feels like diligence, which is asking for sources and then citing them. The step neither took is the cheap one: opening the document. One in 277 is what happens when a whole profession skips a step that takes ninety seconds.
And if 1,163 people representing themselves in court could not tell a real case from an invented one, a second-year with an essay due at nine cannot either. That is not a comment on anybody's intelligence. It is what a well-formed fake looks like.
What these numbers do not establish#
They do not give a fabrication rate for any model, because neither dataset records which tool was used in a way that supports a rate, and the denominator in both cases is documents produced rather than queries made.
They do not show the trend is accelerating in the world rather than in the detection of it. And the court figures cover decisions rather than filings, so they capture the cases a judge chose to rule on, in jurisdictions that publish judgments in a form somebody can compile. Anyone quoting either number as "how often AI makes things up" has changed the claim.
Key sources
- Topaz, C. M. et al. (2026). Fabricated citations: an audit across 2.5 million biomedical papers.
- Charlotin, D. (2026). AI Hallucination Cases Database. Read at source 6 September 2026.
Related SuperSkills research#
What to do about it, the source rule, and for a reading list, how to handle forty readings. On the mechanism, what a hallucination is and why AI sounds so confident. On checking generally, how to know when AI is wrong and the most-quoted AI statistics, checked.
About this research#
Both counts belong to the people who compiled them and are cited above with the dates they were read. Neither is a SuperSkills figure. The reading offered here, that both are floors rather than estimates and that the pro se majority is the number a student should pay attention to, is an interpretation by Rahim Hirji, author of SuperSkills (Kogan Page, 2026), and is marked as an interpretation and not a finding.
Evidence review · SS-2026-189 · Graded against the published rubric
Hirji, R. (2026). How often does AI invent a source?. The SuperSkills evidence base, SS-2026-189. https://thesuperskills.com/research/how-often-does-ai-invent-a-source. Last reviewed 6 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work