← Research
Research · Question

How often is AI wrong?

There is no single rate. The figure depends on the task, who marked it and what counted as wrong.

Last reviewed: 1 October 2026 · Next review due: 1 October 2027

What a news audit, a medical trial and a theory of why models guess say about how often AI is wrong, why a rate means nothing without its task, assessor and rule, and how to measure your own.

Question this page answersAll 1156 questions this research covers

There is no single rate, and any figure quoted without a task attached is misleading. In the largest published audit of AI assistants answering news questions, journalists found at least one significant problem in 45 per cent of answers. In a Nature Medicine trial of medical scenarios, the same kind of model identified the condition in about 95 per cent of cases when tested alone, and people using it identified it in under 35 per cent. Both results are true, and they answer different questions. How often AI is wrong depends on what is asked, who marks the answer and what counts as wrong.

The answer, in one line

It depends on the task. In an EBU and BBC audit of news answers, 45 per cent had at least one significant issue. In a medical trial, models alone found the condition in 94.9 per cent of cases and people using them in under 34.5 per cent.

Share as a card

Definition#

AI error rate: the share of an AI system's outputs, on a stated set of tasks, that a named assessor judges wrong or misleading under a stated marking rule. A rate that omits any of the three, the tasks, the assessor or the rule, cannot be compared with any other. The term is descriptive and is not a SuperSkills coinage.

Share this definition as a card

Forty-five per cent of news answers had a problem, and the definition of a problem decides the figure#

The European Broadcasting Union and the BBC published their study on 22 October 2025. Professional journalists from 22 public service media organisations in 18 countries, working in 14 languages, assessed more than 3,000 responses from ChatGPT, Copilot, Gemini and Perplexity. Each participant put the same 30 core news questions to all four assistants. A significant issue was defined as something that could materially mislead the user, and 45 per cent of answers had at least one. Serious sourcing problems, meaning missing, misleading or incorrect attributions, appeared in 31 per cent, and major accuracy issues, including hallucinated details and outdated information, in 20 per cent. The release reports Gemini worst, at 76 per cent.

The assessors are journalists at the organisations whose own news is the subject of the questions, which gives them expertise and a stake at once. The release counts sourcing and accuracy failures together with other issues, so 45 per cent is not the share of answers that stated a falsehood. Assistants are updated often, which dates the per-assistant figure faster than the headline.

A model that scored 95 per cent alone and a person who scored a third with it#

Bean and colleagues gave 1,298 UK adults ten medical scenarios written by doctors, in a preregistered trial published in Nature Medicine in February 2026. Tested alone, the models identified the condition in 94.9 per cent of cases. Participants using the same models identified a relevant condition in fewer than 34.5 per cent, no better than a control group using whatever they normally would. The error was in the exchange: in 16 of 30 sampled conversations the first message gave only partial information, and the models suggested conditions of which 34.0 per cent were correct. The full account is on should I use AI for medical questions. The point here is that the same model has two error rates, one for a clean question and one for a person who has to ask it.

Models guess because the scoring pays them to#

Kalai, Nachum, Vempala and Zhang argue in a September 2025 preprint that language models produce confident errors because training and evaluation reward guessing over admitting uncertainty. Most dominant benchmarks give an abstention no more credit than a wrong answer, so a model that always guesses outscores one that says when it does not know. Their example is a request for a named person's birthday, with an instruction to answer only if known, to which an open-source model gave three different wrong dates in three attempts. It is a theoretical paper from authors at the firm that builds such models, and as far as a 1 October 2026 search found it has not been peer-reviewed. As an explanation it fits the audit result above: sources invented or attributed wrongly are what a system produces when a plausible answer scores better than silence. Fabricated citations have their own page, how often does AI invent a source.

What none of these figures tell you about your own use#

Fluent wrong answers get caught only by someone who has done the work#

This section is interpretation, kept apart from the evidence above.

A wrong answer in fluent prose looks the same as a right one, and that is the property that makes a rate hard to act on. In Show Your Working (12 July 2026) Rahim Hirji argues that asking a model to show its reasoning does not fix this, because the working is only another answer, as fluent and unaccountable as the first. The person who can catch the slip is the one who has done that kind of work by hand and knows where the middle tends to go wrong. That claim is an argument and no study here tests it, though Bean's result, where people could not turn a strong model into a good decision, sits comfortably beside it. The capability involved is the subject of how do I know when AI is wrong.

Measuring your own error rate on twenty real tasks#

Key sources

On checking, how do I know when AI is wrong. On confident delivery, why does AI sound so confident. On invented references, how often does AI invent a source. On the medical case, should I use AI for medical questions.

About this research#

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. The audit release and toolkit and the Kalai preprint were read at source on 1 October 2026, and the Bean figures were read on 30 September 2026; all are graded in the evidence base. AI error rate is a descriptive term and not a SuperSkills coinage.

How this research works  ·  Reviewed quarterly  ·  Found an error? Tell me and it is corrected on the page.

Evidence review · SS-2026-376 · Graded against the published rubric · 1 peer-reviewed study, 1 working paper and 1 interested audit

Cite this page

Hirji, R. (2026). How often is AI wrong?. The SuperSkills evidence base, SS-2026-376. https://thesuperskills.com/research/how-often-is-ai-wrong. Last reviewed 1 October 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

How often is AI wrong?

It depends on the task. In an EBU and BBC audit of news answers, 45 per cent had at least one significant issue. In a medical trial, models alone found the condition in 94.9 per cent of cases and people using them in under 34.5 per cent. There is no single rate.

How accurate are AI assistants on news?

In the audit published on 22 October 2025, journalists found a significant issue in 45 per cent of more than 3,000 answers from four assistants, serious sourcing problems in 31 per cent and major accuracy issues in 20 per cent. The assessors have a stake in the result.

Why does AI give wrong answers so confidently?

One argument, in an unrefereed 2025 preprint by Kalai and colleagues, is that training and evaluation reward guessing over admitting uncertainty, because most benchmarks score a refusal no higher than an error. It is an explanation and not a measured rate.

Is AI more accurate than people?

Not as a general claim. The Nature Medicine trial found models alone scored far higher than people who used them, so the answer depends on whether the comparison is the model by itself or the model in a person's hands.

How do I find out how often AI is wrong at my work?

Sample twenty real tasks, have someone who knows the answer mark the output, and note which errors looked plausible. Twenty is too few for a precise rate and enough to show whether errors are rare, common or clustered.

In this hub

Judgement, oversight and accountability

Who decides, who checks, and who is answerable when the machine was involved.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory  ·  Boards  ·  Enquire