You cannot tell from the output. That is the whole problem, and every technique that promises otherwise is selling you something. A language model's confidence is a property of its writing style rather than of its knowledge, so fluency, structure, hedging and citation all look identical whether the answer is right or wrong. The only reliable signal is external to the text: knowing, in advance and for your own domain, which categories of question the model handles well and which it handles badly. That knowledge is local, it takes months to build, it does not transfer between fields, and almost nobody has it. Which is also why it is one of the few capabilities that does not commoditise.
Why the surface tells you nothing
The most important result here is the jagged technological frontier. In 2023 Dell'Acqua and colleagues, with Boston Consulting Group and researchers at Harvard, MIT and Wharton, gave 758 consultants access to GPT-4. Inside the model's competence they were dramatically better and faster. On a task deliberately placed just outside it, consultants using AI performed worse than consultants using none.
The word doing the work is jagged. The frontier is not a smooth boundary where performance degrades gracefully as questions get harder. Competence on one task tells you almost nothing about competence on an adjacent one, and the model's manner does not change as it crosses over. Those consultants were not careless. They were reading output that gave them no signal.
Nor can you rely on explanation to save you. Dzindolet and colleagues found in 2003 that explaining why an automated aid might err increased reliance on it, even when the restored trust was unwarranted. Being told how a system can fail can make you trust it more. That is a genuinely awkward finding for anyone whose safeguard is a disclaimer.
And there is a well-documented case of the verification trap closing completely. In January 2025 the High Court in Pietermaritzburg dealt with counsel who had cited authorities that did not exist. The judge tested one citation by asking ChatGPT, which falsely confirmed the case was real. The verification method was the same class of system that produced the error.
Where the evidence is uncertain
Model capability moves quickly, and the jagged-frontier experiment used GPT-4 in 2023. The specific tasks that sat outside the frontier then may sit comfortably inside it now. What has not changed, and shows no sign of changing, is that the boundary remains jagged and remains invisible from the output. A better model moves the line without drawing it.
There is also an honest limit on the advice below. Building a frontier map requires enough domain expertise to recognise a wrong answer in the first place, which means it is available to experienced practitioners and largely unavailable to anyone early in a career. That is not a gap this page can close, and it is precisely the trap described in synthetic seniority.
The SuperSkills view
Reframe the question. "How do I know when AI is wrong?" invites a search for tells, and there are none. The useful question is "in what circumstances is this system likely to be wrong for the kind of work I do?" That is answerable, it is specific to you, and it is the actual skill.
There are recognisable categories where the risk is elevated, and they are worth learning as a set. Anything requiring a precise fact that is rare, recent or contested. Anything where the correct answer depends on context the model was never given, which includes most decisions inside an organisation. Anything at the edge of a domain rather than its centre, where training data thins. Anything where the plausible answer and the correct answer differ, which is the most dangerous class of all, because plausibility is exactly what the system optimises. And anything where you would not be able to detect the error yourself, which is the honest test.
That last one is the point at which this connects to everything else on this site. Verification is not a separate activity from expertise; it is expertise, applied. You cannot check an answer in a domain where you never built competence, which means every repetition handed to the machine is also a small reduction in your ability to supervise the machine. That is capability debt in its most immediate form, and it is why I argue that verification should be resourced and paid as skilled work rather than treated as administrative residue.
Building your own frontier map
This is the single highest-return habit available, and it costs a few minutes a week.
- Keep a running note of every confident wrong answer you catch in your own domain. Not the amusing ones, the plausible ones. After six months you will have something no competitor can copy, because it is built from your work.
- Write your own answer before you prompt, even briefly. You cannot notice a divergence from a position you never held. This is the individual form of Human at the Start.
- Verify against a different kind of source, never against another model. Primary documents, a person who knows, the original study. The South African case is the cautionary version of getting this wrong.
- Set the rejection criteria before you see the output. Deciding what would make you say no is much easier before an answer is sitting in front of you looking finished.
- Treat unusual confidence on an unusual question as a signal. Not proof of error, but the moment to slow down.
Related SuperSkills research
The underlying tendency is automation bias. On why verification is undervalued, the verifier's discount. On decision design, human and AI decision making. On why this capability appreciates, staying valuable in the age of AI and why "learn to prompt" is weak career advice.
Key research and primary sources
- Dell'Acqua, F. et al. (2023). Navigating the Jagged Technological Frontier. Harvard Business School and BCG working paper.
- Dzindolet, M. T. et al. (2003). The role of trust in automation reliance. International Journal of Human-Computer Studies, 58(6).
- Parasuraman, R. and Manzey, D. H. (2010). Complacency and Bias in Human Use of Automation. Human Factors, 52(3).
- Yu, F. et al. (2024). Heterogeneity and predictors of the effects of AI assistance on radiologists. Nature Medicine, 30(3).
- Mavundla v MEC: COGTA KwaZulu-Natal [2025] ZAKZPHC 2. High Court of South Africa, Pietermaritzburg.
About this research
Rahim Hirji is the author of SuperSkills: The Seven Human Skills for the Age of AI (Kogan Page, 2026) and the founder of The SuperSkills Intelligence Company. The jagged technological frontier is Dell'Acqua and colleagues' term, not his. Findings are attributed to the studies that produced them and kept separate from the interpretation. Given how quickly model capability moves, this page is on a 90-day review cycle.
Cite this
Hirji, R. (2026). How do I know when AI is wrong? The SuperSkills Intelligence Company. Last reviewed 26 August 2026. thesuperskills.com/research/how-do-i-know-when-ai-is-wrong