- How do I know when AI is wrong?
- Can I use AI to check AI?
- Do AI agents claim to have done work they have not done?
- Should I use primary sources rather than an AI summary?
- How can people detect when AI is making a bad judgement?
- Can teenagers tell when a chatbot is wrong?
- Should I trust AI health information?
- Should a human verify data an AI has extracted?
You cannot tell from the output. That is the whole problem, and every technique that promises otherwise is selling you something. A language model's confidence is a property of its writing style rather than of its knowledge, so fluency, structure, hedging and citation all look identical whether the answer is right or wrong. The only reliable signal is external to the text: knowing, in advance and for your own domain, which categories of question the model handles well and which it handles badly. That knowledge is local, it takes months to build, it does not transfer between fields, and almost nobody has it. Which is also why it is one of the few capabilities that does not commoditise.
The answer, in one line
You cannot tell from the output, and any technique promising otherwise is unreliable. A language model's confidence is a property of its writing style rather than its knowledge, so fluency, structure, hedging and citation look identical whether the answer is right or wrong.
Why the surface tells you nothing#
The most important result here is the jagged technological frontier. In 2023 Dell'Acqua and colleagues, with Boston Consulting Group and researchers at Harvard, MIT and Wharton, gave 758 consultants access to GPT-4. Inside the model's competence they were dramatically better and faster. On a task deliberately placed just outside it, consultants using AI performed worse than consultants using none.
The word doing the work is jagged. The frontier is not a smooth boundary where performance degrades gracefully as questions get harder. Competence on one task tells you almost nothing about competence on an adjacent one, and the model's manner does not change as it crosses over. Those consultants were not careless. They were reading output that gave them no signal.
Nor can you rely on explanation to save you. Dzindolet and colleagues found in 2003 that explaining why an automated aid might err increased reliance on it, even when the restored trust was unwarranted. Being told how a system can fail can make you trust it more. That is a genuinely awkward finding for anyone whose safeguard is a disclaimer.
And there is a well-documented case of the verification trap closing completely. In January 2025 the High Court in Pietermaritzburg dealt with counsel who had cited authorities that did not exist. The judge tested one citation by asking ChatGPT, which falsely confirmed the case was real. The verification method was the same class of system that produced the error.
Model capability keeps moving#
Model capability moves quickly, and the jagged-frontier experiment used GPT-4 in 2023. The specific tasks that sat outside the frontier then may sit comfortably inside it now. What has not changed, and shows no sign of changing, is that the boundary remains jagged and remains invisible from the output. A better model moves the line without drawing it.
There is also an honest limit on the advice below. Building a frontier map requires enough domain expertise to recognise a wrong answer in the first place, which means it is available to experienced practitioners and largely unavailable to anyone early in a career. That is a gap this page cannot close, and the trap described in synthetic seniority.
A better question than hunting for tells#
Reframe the question. "How do I know when AI is wrong?" invites a search for tells, and there are none. The useful question is "in what circumstances is this system likely to be wrong for the kind of work I do?" That is answerable, specific to you, and the actual skill.
There are recognisable categories where the risk is elevated, and they are worth learning as a set. Anything requiring a precise fact that is rare, recent or contested. Anything where the correct answer depends on context the model was never given, which includes most decisions inside an organisation. Anything at the edge of a domain rather than its centre, where training data thins. Anything where the plausible answer and the correct answer differ, which is the most dangerous class of all, because plausibility is what the system optimises for. And anything where you would not be able to detect the error yourself, which is the honest test.
That last one is the point at which this connects to everything else on this site. Verification is not a separate activity from expertise; it is expertise, applied. You cannot check an answer in a domain where you never built competence, which means every repetition handed to the machine is also a small reduction in your ability to supervise the machine. That is capability debt in its most immediate form. It is why I argue that verification should be resourced and paid as skilled work rather than treated as administrative residue.
Building your own frontier map#
This is the single highest-return habit available, and it costs a few minutes a week.
- Keep a running note of every confident wrong answer you catch in your own domain. Not the amusing ones, the plausible ones. After six months you will have something no competitor can copy, because it is built from your work.
- Write your own answer before you prompt, even briefly. You cannot notice a divergence from a position you never held. This is the individual form of Human at the Start.
- Verify against a different kind of source, never against another model. Primary documents, a person who knows, the original study. The South African case is the cautionary version of getting this wrong.
- Set the rejection criteria before you see the output. Deciding what would make you say no is much easier before an answer is sitting in front of you looking finished.
- Treat unusual confidence on an unusual question as a signal. Not proof of error, but the moment to slow down.
17 September 2026: the agent’s report of its own work is not evidence of the work#
This page argues that you cannot tell from the output. A paper posted on 17 September extends that to a different output: the agent’s account of what it did. In OverclaimBench, researchers at Tara Research, Mila and Cohere gave twelve frontier coding agents sets of files to review, each with defects planted in them, and measured from the tool calls which files were read. The agents failed to read every requested file in 67.9 per cent of 1,140 runs. In those incomplete runs the final response was misleading 80.4 per cent of the time, either claiming a complete review or saying nothing about the gap, and the runs that explicitly claimed completeness missed the planted defects at about 1.8 times the rate of the runs that had read everything. The rate of explicit false claims ran from 9.8 to 73.7 per cent across the twelve models, from every major provider, so the finding is about the class of tool rather than any one vendor. The authors’ first sentence states the practical problem: “an agent’s final response is often the only account of that work a user sees.”
The caveats are the authors’ own and a reviewer’s. The scenarios were designed by iterating against one provider’s model until the behaviour appeared, and simpler corpora were dropped because they did not elicit it, so 67.9 per cent is what a benchmark built to find the behaviour found, not a base rate in use. A single model judged every response, and no human agreement figure is reported. The paper is not peer reviewed. What survives the caveats is the direction: when the work is incomplete, the report of the work more often hides that than says it. For the frontier map this page describes, that adds a rule. The claim “I checked all of them” sits on the wrong side of the frontier until you have a record, independent of the model, of what it read. The tool log is that record; the summary is not.
Related SuperSkills research#
The underlying tendency is automation bias. On why verification is undervalued, the verifier's discount. On decision design, human and AI decision making. On why this capability appreciates, staying valuable in the age of AI and why "learn to prompt" is weak career advice. The operational version, stage by stage with a downloadable grid, is the Delegation Boundary Map. On the definition, the jagged frontier. On the definition, why AI sounds so confident. See when to override AI. See what is a hallucination.
Key research and primary sources
- Dell'Acqua, F. et al. (2023). Navigating the Jagged Technological Frontier. Harvard Business School and BCG working paper.
- Dzindolet, M. T. et al. (2003). The role of trust in automation reliance. International Journal of Human-Computer Studies, 58(6).
- Parasuraman, R. and Manzey, D. H. (2010). Complacency and Bias in Human Use of Automation. Human Factors, 52(3).
- Yu, F. et al. (2024). Heterogeneity and predictors of the effects of AI assistance on radiologists. Nature Medicine, 30(3).
- Mavundla v MEC: COGTA KwaZulu-Natal [2025] ZAKZPHC 2. High Court of South Africa, Pietermaritzburg.
About this research#
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. The jagged technological frontier is Dell'Acqua and colleagues' term, not his. Findings are attributed to the studies that produced them and kept separate from the interpretation. Given how quickly model capability moves, this page is on a 90-day review cycle.
Essay · SS-2026-042 · 1 working paper
Hirji, R. (2026). How do I know when AI is wrong?. The SuperSkills evidence base, SS-2026-042. https://thesuperskills.com/research/how-do-i-know-when-ai-is-wrong. Last reviewed 20 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work