Because confidence is a property of the writing, not of the knowledge. A language model produces the most probable continuation of a text, and the most probable continuation of a well-formed question is a well-formed answer. Hedging, qualification and expressions of doubt are themselves stylistic patterns the model can produce, so they appear where the training text would have contained them rather than where the model is actually uncertain. Fluency and accuracy are generated by the same process and are not connected.
The short version
The system is not reporting how sure it is. It is producing text that resembles text written by someone who was sure. Those are completely different things, and only one of them is visible to you.
Three reasons the confidence is so convincing
It has no signal to give you. A person who half-remembers something usually sounds like a person who half-remembers something. That correlation between internal uncertainty and outward manner is what we rely on in daily life, and it is absent here. Models do carry internal probability distributions, but what surfaces in the prose is a style, not a calibrated reading of it.
Training rewards helpfulness. Systems tuned on human preference data learn what people rate highly, and people rate confident, complete, well-structured answers above hesitant ones. Refusing or hedging is penalised more visibly than being wrong, because the wrongness is often not detected in the moment.
Errors arrive in the correct format. A fabricated citation has authors, a year, a journal and a volume number. A wrong figure has the right number of digits. The form is right even when the content is not, and form is what a reader scans first.
What this does to the person reading
It interacts badly with a documented tendency. Parasuraman and Manzey's review found automation bias and complacency appear in experts as well as novices, resist training and worsen under workload. Logg and colleagues found people often weight algorithmic advice more heavily than human advice, with domain experts the notable exception.
And awareness is a weaker defence than it feels. Dzindolet and colleagues found that explaining how an automated aid can fail could increase reliance on it. Being told the system might be wrong does not reliably make people check.
The clearest illustration is a court record. In January 2025 the High Court in Pietermaritzburg dealt with counsel who had cited authorities that did not exist. The judge tested one citation by asking ChatGPT, which confirmed the case was real and then confirmed it addressed a point it could not have addressed. Confident, formatted, and wrong twice.
So how do you tell?
Not from the text. That is the honest answer and it is why the useful question is different: in what circumstances is this system likely to be wrong for the work I do? Elevated-risk categories include anything requiring a precise fact that is rare, recent or contested; anything depending on context the model was never given; anything at the edge of a domain rather than its centre; and anything where the plausible answer and the correct answer differ, which is the worst class because plausibility is what the system optimises.
The practical version, including how to build a map of your own domain's failure patterns, is in how do I know when AI is wrong.
What is uncertain
Calibration is an active research area and some systems do expose uncertainty estimates, though rarely in the interfaces most people use. Whether future systems will communicate doubt in a way that is both accurate and actually attended to is open. Treat this page as describing systems as they behave in 2026, not as a permanent property of the technology.
Related SuperSkills research
On the boundary that produces the errors, the jagged frontier. On the tendency it exploits, automation bias. On who is supposed to catch it, who owns verification. On why review at the end is the weakest control, human in the loop is not a safeguard.
Key sources
- Parasuraman, R. and Manzey, D. H. (2010). Complacency and Bias in Human Use of Automation. Human Factors, 52(3).
- Logg, J. M. et al. (2019). Algorithm appreciation. Organizational Behavior and Human Decision Processes, 151.
- Dzindolet, M. T. et al. (2003). The role of trust in automation reliance. IJHCS, 58(6).
- Mavundla v MEC: COGTA KwaZulu-Natal [2025] ZAKZPHC 2. High Court of South Africa, Pietermaritzburg.
About this page
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026) and founder of The SuperSkills Intelligence Company. This describes established behaviour of language models and is not a SuperSkills coinage or claim. On a 90-day review cycle.
Cite this
Hirji, R. (2026). Why does AI sound so confident? The SuperSkills Intelligence Company. Last reviewed 26 August 2026. thesuperskills.com/research/why-does-ai-sound-so-confident