← Research
Research

How often does AI invent a source?

Both counts come from professions built on checking references. Both are floors rather than estimates.

Last reviewed: 6 September 2026

One in 277 on PubMed, 2,022 court decisions and who actually filed them, why both numbers are floors, and the property that makes a fabricated citation different from ordinary error.

Question this page answersQuestion this page partly answersAll 811 questions this research covers

Often enough that it now has two counts. In published, peer-reviewed medicine, one paper in 277 on PubMed cited a study that does not exist in early 2026, against one in 2,828 in 2023. In court, 2,022 decisions worldwide have recorded somebody filing material that was never written. Both counts come from people whose job is checking sources, working without a deadline at two in the morning, and both are floors rather than estimates.

The answer, in one line

Two counted floors exist.

Share as a card

Definition#

A fabricated citation: a reference produced in the correct form, with a plausible author, a real journal and a fitting year, for a work that was never published. It is distinct from a misquotation, because there is no document to have misquoted.

Share this definition as a card

One paper in 277#

An audit across 2.5 million biomedical papers found 4,046 fabricated references across 2,810 papers. The rate of papers carrying at least one rose from 1 in 2,828 in 2023, to 1 in 458 in 2025, to 1 in 277 in the first seven weeks of 2026: roughly twelvefold in three years.

Review articles ran 57 per cent higher than other paper types, which makes sense and is the part that should worry a student most. A review is where somebody summarises a literature they have not personally run, which is structurally the same act as asking a model what a field says.

These are papers that went through peer review, written by researchers, in journals indexed on PubMed.

Two thousand court decisions, and most of them are not lawyers#

The AI Hallucination Cases database tracks legal decisions in which a court found that generative AI had produced hallucinated content in material put before it. Read at source on 6 September 2026, it holds 2,022 decisions.

By jurisdiction: the United States 1,379, Canada 217, Australia 110, the United Kingdom 69, Israel 57, Brazil 41, India and Italy 15 each, France 13, Germany 11, and roughly thirty further countries.

The figure most reporting leaves out is who filed the material. Pro se litigants, people representing themselves, account for 1,163 of the recorded parties. Lawyers account for 805. Judges appear 31 times, experts 15, prosecutors 5.

Read those two numbers together. The professional half of that count is people with a duty to check, insurance to lose and a regulator watching. The larger half is people with none of those things, doing what a student does: reaching for a tool under pressure, in an unfamiliar system, with no way of telling a real citation from an invented one.

Why both numbers are floors#

Neither count measures how often models fabricate. Both measure how often somebody was caught.

The court database says so in its own words: it tracks decisions and "does not track the (necessarily wider) universe of all fake citations or use of AI in court filings". A fabricated citation that nobody checks is absent by construction. The same applies to the PubMed audit, which finds what an automated check can find in a published record.

Growth in either count also mixes three things: more fabrication, better detection, and better compilation. The court database is updated daily and its author calls it a work in progress, so any figure taken from it has to carry the date it was read. This page read it on 6 September 2026 and the number will be higher by the time you do.

The property that makes this different from ordinary error#

A wrong citation has always been possible. What is new is the form the wrongness takes.

A model does not produce an obviously broken reference. It produces one in exactly the shape a real citation takes, because that shape is what it learned. Plausible author, real journal, year that fits the argument, page range in the right format. Every instinct a reader has for spotting a weak source was trained on sources that exist, and none of those instincts fires.

That is why the first move of the source rule is to establish that the source exists, before any judgement about whether it is good. Caulfield's SIFT and the CRAAP test both begin at the second question, because both were written when the first one did not need asking.

What a student should take from two numbers about doctors and lawyers#

Not that AI is untrustworthy, which is too broad to act on. Something narrower and more useful: the people in these counts were not careless, and being careful is not the defence.

A researcher writing a review and a litigant filing a brief both did the thing that feels like diligence, which is asking for sources and then citing them. The step neither took is the cheap one: opening the document. One in 277 is what happens when a whole profession skips a step that takes ninety seconds.

And if 1,163 people representing themselves in court could not tell a real case from an invented one, a second-year with an essay due at nine cannot either. That is not a comment on anybody's intelligence. It is what a well-formed fake looks like.

What these numbers do not establish#

They do not give a fabrication rate for any model, because neither dataset records which tool was used in a way that supports a rate, and the denominator in both cases is documents produced rather than queries made.

They do not show the trend is accelerating in the world rather than in the detection of it. And the court figures cover decisions rather than filings, so they capture the cases a judge chose to rule on, in jurisdictions that publish judgments in a form somebody can compile. Anyone quoting either number as "how often AI makes things up" has changed the claim.

Key sources

What to do about it, the source rule, and for a reading list, how to handle forty readings. On the mechanism, what a hallucination is and why AI sounds so confident. On checking generally, how to know when AI is wrong and the most-quoted AI statistics, checked.

About this research#

Both counts belong to the people who compiled them and are cited above with the dates they were read. Neither is a SuperSkills figure. The reading offered here, that both are floors rather than estimates and that the pro se majority is the number a student should pay attention to, is an interpretation by Rahim Hirji, author of SuperSkills (Kogan Page, 2026), and is marked as an interpretation and not a finding.

Evidence review · SS-2026-189 · Graded against the published rubric

Cite this page

Hirji, R. (2026). How often does AI invent a source?. The SuperSkills evidence base, SS-2026-189. https://thesuperskills.com/research/how-often-does-ai-invent-a-source. Last reviewed 6 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

How often does AI make up citations?

Two counted floors exist. In published biomedical research, an audit of 2.5 million papers found the rate of papers carrying at least one fabricated reference rising from 1 in 2,828 in 2023 to 1 in 458 in 2025 and 1 in 277 in the first seven weeks of 2026, with review articles running 57 per cent higher than other paper types. In law, a database of court decisions recorded 2,022 cases worldwide as at 6 September 2026. Neither measures how often models fabricate; both measure how often somebody was caught.

Is it only lawyers filing fake AI citations in court?

No, and this is the figure most reporting leaves out. In the AI Hallucination Cases database, pro se litigants, meaning people representing themselves, account for 1,163 of the recorded parties against 805 lawyers. Judges themselves appear 31 times, experts 15 and prosecutors 5. By jurisdiction the United States has 1,379, Canada 217, Australia 110 and the United Kingdom 69, with roughly thirty further countries represented. The professional half has a duty to check and insurance to lose; the larger half has neither.

Why are fabricated AI citations so hard to spot?

Because a model does not produce an obviously broken reference. It produces one in exactly the shape a real citation takes, since that shape is what it learned: plausible author, real journal, a year that fits the argument, a page range in the right format. Every instinct a reader has for spotting a weak source was trained on sources that exist, so none of those instincts fires. That is why the first move has to be establishing that the source exists, before any judgement about whether it is good.

Are these numbers getting worse?

The counts are rising and that is not the same claim. Growth in either mixes three things: more fabrication, better detection, and better compilation. The court database is updated daily and its author describes it as a work in progress, so any figure taken from it must carry the date it was read. It also states that it tracks decisions and does not track the wider universe of all fake citations or AI use in court filings, so undetected instances are absent by construction.

In this hub

For students

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory for CEOs and boards  ·  Enquire