There is no single rate, and any figure quoted without a task attached is misleading. In the largest published audit of AI assistants answering news questions, journalists found at least one significant problem in 45 per cent of answers. In a Nature Medicine trial of medical scenarios, the same kind of model identified the condition in about 95 per cent of cases when tested alone, and people using it identified it in under 35 per cent. Both results are true, and they answer different questions. How often AI is wrong depends on what is asked, who marks the answer and what counts as wrong.
The answer, in one line
It depends on the task. In an EBU and BBC audit of news answers, 45 per cent had at least one significant issue. In a medical trial, models alone found the condition in 94.9 per cent of cases and people using them in under 34.5 per cent.
Definition#
AI error rate: the share of an AI system's outputs, on a stated set of tasks, that a named assessor judges wrong or misleading under a stated marking rule. A rate that omits any of the three, the tasks, the assessor or the rule, cannot be compared with any other. The term is descriptive and is not a SuperSkills coinage.
Forty-five per cent of news answers had a problem, and the definition of a problem decides the figure#
The European Broadcasting Union and the BBC published their study on 22 October 2025. Professional journalists from 22 public service media organisations in 18 countries, working in 14 languages, assessed more than 3,000 responses from ChatGPT, Copilot, Gemini and Perplexity. Each participant put the same 30 core news questions to all four assistants. A significant issue was defined as something that could materially mislead the user, and 45 per cent of answers had at least one. Serious sourcing problems, meaning missing, misleading or incorrect attributions, appeared in 31 per cent, and major accuracy issues, including hallucinated details and outdated information, in 20 per cent. The release reports Gemini worst, at 76 per cent.
The assessors are journalists at the organisations whose own news is the subject of the questions, which gives them expertise and a stake at once. The release counts sourcing and accuracy failures together with other issues, so 45 per cent is not the share of answers that stated a falsehood. Assistants are updated often, which dates the per-assistant figure faster than the headline.
A model that scored 95 per cent alone and a person who scored a third with it#
Bean and colleagues gave 1,298 UK adults ten medical scenarios written by doctors, in a preregistered trial published in Nature Medicine in February 2026. Tested alone, the models identified the condition in 94.9 per cent of cases. Participants using the same models identified a relevant condition in fewer than 34.5 per cent, no better than a control group using whatever they normally would. The error was in the exchange: in 16 of 30 sampled conversations the first message gave only partial information, and the models suggested conditions of which 34.0 per cent were correct. The full account is on should I use AI for medical questions. The point here is that the same model has two error rates, one for a clean question and one for a person who has to ask it.
Models guess because the scoring pays them to#
Kalai, Nachum, Vempala and Zhang argue in a September 2025 preprint that language models produce confident errors because training and evaluation reward guessing over admitting uncertainty. Most dominant benchmarks give an abstention no more credit than a wrong answer, so a model that always guesses outscores one that says when it does not know. Their example is a request for a named person's birthday, with an instruction to answer only if known, to which an open-source model gave three different wrong dates in three attempts. It is a theoretical paper from authors at the firm that builds such models, and as far as a 1 October 2026 search found it has not been peer-reviewed. As an explanation it fits the audit result above: sources invented or attributed wrongly are what a system produces when a plausible answer scores better than silence. Fabricated citations have their own page, how often does AI invent a source.
What none of these figures tell you about your own use#
- Different tasks. News questions and medical vignettes say nothing about contract summaries, code or arithmetic. No figure here transfers across them.
- Different versions. Both studies tested models available in 2025, and assistants change often. The trial's authors call their figures a lower bound for newer models.
- Different markers. A journalist, a doctor and a benchmark mark differently. The 94.9 per cent is a model against an answer key and the 45 per cent is a human reading for what could mislead.
- No rate for the case where you already know the answer. Every study above counts errors that someone else found. How many go uncaught inside a firm is not measured.
Fluent wrong answers get caught only by someone who has done the work#
This section is interpretation, kept apart from the evidence above.
A wrong answer in fluent prose looks the same as a right one, and that is the property that makes a rate hard to act on. In Show Your Working (12 July 2026) Rahim Hirji argues that asking a model to show its reasoning does not fix this, because the working is only another answer, as fluent and unaccountable as the first. The person who can catch the slip is the one who has done that kind of work by hand and knows where the middle tends to go wrong. That claim is an argument and no study here tests it, though Bean's result, where people could not turn a strong model into a good decision, sits comfortably beside it. The capability involved is the subject of how do I know when AI is wrong.
Measuring your own error rate on twenty real tasks#
- Ask for the rate on your task. When a vendor or a colleague quotes an accuracy figure, ask what the tasks were, who marked them and what counted as wrong.
- Sample your own work. Take twenty real requests, have someone who knows the answer mark the output, and note which errors were plausible. Twenty is too few for a precise rate and enough to show whether the errors are rare, common or clustered.
- Set checking by the cost of a miss. Spend most scrutiny where a wrong answer is expensive or goes unseen, such as numbers, sources, dates and anything sent to a customer.
- Reward a stated doubt. If the scoring that trains models pays for guessing, a team that rewards "I could not verify this" is correcting for it. That is an inference from the Kalai argument and not a tested practice.
Key sources
- European Broadcasting Union and BBC (2025). Largest study of its kind shows AI assistants misrepresent news content 45% of the time. EBU release, 22 October 2025, with the News Integrity in AI Assistants toolkit. Graded entry.
- Bean, A. M. et al. (2026). Reliability of LLMs as medical assistants for the general public: a randomized preregistered study. Nature Medicine 32(2), 609 to 615. Graded entry.
- Kalai, A. T., Nachum, O., Vempala, S. S. and Zhang, E. (2025). Why Language Models Hallucinate. arXiv:2509.04664, preprint. Graded entry.
Related SuperSkills research#
On checking, how do I know when AI is wrong. On confident delivery, why does AI sound so confident. On invented references, how often does AI invent a source. On the medical case, should I use AI for medical questions.
About this research#
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. The audit release and toolkit and the Kalai preprint were read at source on 1 October 2026, and the Bean figures were read on 30 September 2026; all are graded in the evidence base. AI error rate is a descriptive term and not a SuperSkills coinage.
Evidence review · SS-2026-376 · Graded against the published rubric · 1 peer-reviewed study, 1 working paper and 1 interested audit
Hirji, R. (2026). How often is AI wrong?. The SuperSkills evidence base, SS-2026-376. https://thesuperskills.com/research/how-often-is-ai-wrong. Last reviewed 1 October 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work