Two problems are usually treated as one, and they need different fixes. The first is that models are badly calibrated when asked to state confidence in words, with expected calibration error above 0.37 for four of five systems tested and stated confidences clustering between 80 and 100 per cent. The second is that even a perfectly calibrated word is read differently by every reader: given text using the IPCC's probability vocabulary, people put "very likely" at 62 per cent where the guidelines mean above 90. So the answer is a fixed vocabulary published with numerical ranges, expressed in the first person, separating how likely the claim is from how good the basis for it is, and with silence forbidden rather than treated as neutral. Three of those four have experimental evidence behind them; the fourth is a tradecraft rule that a national intelligence service has run for a decade. Almost none of it is deployed.
The answer, in one line
With a fixed vocabulary whose terms are published with numerical ranges, in the first person, separating how likely the claim is from how good the basis for it is, and with the absence of a marker forbidden rather than treated as neutral.
The problem was solved in 1964 and the solution never took#
Sherman Kent ran the Board of National Estimates at the CIA. In an essay published in Studies in Intelligence in 1964, he describes the moment the problem became visible to him. A policymaker asked what the phrase "serious possibility" had meant in a national estimate. Kent replied that his own view was around 65 to 35 in favour. The reply jolted the man, who had read it as far lower odds. So Kent asked his colleagues on the Board, all of whom had signed the same sentence.
It was another jolt to find that each Board member had had somewhat different odds in mind and the low man was thinking of about 20 to 80, the high of 80 to 20. The rest ranged in between.
The authors of a single agreed sentence disagreed with each other by a factor of sixteen about what it meant. Kent's response was a chart assigning percentage bands to a controlled vocabulary, and one absolute rule: "The word 'possible' (and its cognates) must not be modified." His verdict on how it was received, two paragraphs from the end: "We are in disarray."
Kent's essay is an essay, not a study. The spread he reports is an informal poll of a handful of colleagues and the number of them is not stated. What makes it worth the space is what happened next. The American intelligence community eventually did adopt his solution, and still runs it. Intelligence Community Directive 203 requires that expressions of likelihood use one of two named rows of seven terms, published against percentage bands: 01-05, 05-20, 20-45, 45-55, 55-80, 80-95, 95-99, running from "almost no chance" to "almost certain". Analysts are told not to mix rows, and a product that mixes them must carry a disclaimer.
The directive also contains a rule that no AI interface this research has seen observes, the most transferable thing in the whole document. Likelihood and confidence are different quantities and may not be combined: a product expressing confidence in an assessment "must not combine a confidence level and a degree of likelihood, which refers to an event or development, in the same sentence". How probable the thing is, and how good the basis for saying so is, are two axes. Systems collapse them constantly, and a user reading "high confidence" cannot tell which one they have been given.
The climate scientists ran the experiment, and handing readers the key did not help#
The IPCC faced Kent's problem at scale and published a translation table: likely means above 66 per cent, very likely above 90, and so on. Budescu, Por and Broomell tested whether it works, on a nationally representative US panel. Of 841 invited panel members, 556 completed the survey, a 66 per cent response, split across a control condition, a condition where the translation table was supplied, and one where numerical ranges appeared alongside the words in the text itself.
The results are worse than the pessimistic reading. Mean estimates for the four terms tested were 41 for "very unlikely", 44 for "unlikely", 54 for "likely" and 62 for "very likely". Four distinct terms spanning the whole probability range were read as four numbers clustered within twenty-one points of even odds. Consistency with the guidelines ran at 20.76 per cent in the control, 18.81 per cent when the translation table was supplied, and 30.12 per cent when numbers sat beside the words. Twenty-four per cent of respondents produced no response consistent with the guidelines at all, and only 6 per cent produced six or more.
Read the middle number again. Giving readers the key made them numerically worse than giving them nothing. Only embedding the range in the sentence they were reading helped, and it lifted consistency to under a third. The authors describe the pattern as regressive: "the IPCC intends to convey a probability of at least 0.90, but the typical respondent (i.e., the median response) interprets this to mean only about 0.65-0.75". A follow-up across 25 samples in 24 countries and 17 languages found the same shape, reporting that "laypeople interpret IPCC statements as conveying probabilities closer to 50% than intended by the IPCC authors" and that the qualitative patterns are "remarkably stable across all samples and languages".
The design implication is direct and slightly humbling. Publishing a glossary of what your system's hedging words mean will not work. The number has to be in the sentence.
The model's number is better than the reader's, and neither is good#
Steyvers and colleagues put the two halves together in Nature Machine Intelligence, with 301 participants across two experiments answering questions with model help. They measure how well a signal separates correct answers from incorrect ones, and compare the model's own confidence against what a reader takes from the model's explanation.
The model's internal confidence discriminates reasonably: AUC 0.751 for GPT-3.5 and 0.746 for PaLM2 on multiple choice, 0.781 for GPT-4o on short answer. Participants reading the default explanations came in at 0.589, 0.602 and 0.592, which they describe as "only slightly better than random guessing". They name the two shortfalls the calibration gap and the discrimination gap, and attribute the first to the reader rather than the machine: "human miscalibration is primarily due to overconfidence, indicating that people generally believe that LLMs are more accurate than they actually are."
The finding with the most direct product consequence concerns length. "Long explanations led to significantly higher confidence than the short explanations", and yet the additional information "did not enable participants to better discriminate between probably correct and incorrect answers", with mean participant AUC at 0.54 for long explanations. More detail made readers surer without making them righter. Any interface whose answer to a hard question is to say more is doing this.
On the machine side, Xiong and colleagues tested five models across eight datasets and report average expected calibration error for plain verbalised confidence of 0.520 for GPT-3, 0.461 for Vicuna, 0.436 for LLaMA 2, 0.377 for GPT-3.5 and 0.180 for GPT-4. Even GPT-4's average AUROC of 62.7 per cent sits close to the 50 per cent chance line. Their observation about the shape of the numbers is the memorable one: stated confidences arrive "as multiples of 5 and with most values ranging between the 80% to 100% range", which they suggest means the models "might be imitating human expressions when verbalizing confidence". A model asked how sure it is produces a plausible human-sounding percentage, and that is a different act from measuring anything. This is why AI sounds so confident, restated as a calibration statistic.
Hedging works, and it costs something the product manager will notice#
Kim, Liao, Vorvoreanu, Ballard and Wortman Vaughan ran the experiment that matters commercially. Eight yes-or-no medical questions, a system whose answers were correct exactly half the time, and three versions of identical content: plain, hedged in the first person ("I'm not certain, but it seems to me"), and hedged impersonally ("It's unclear, but it seems like"). Of 656 responses collected, 252 were excluded on pre-registered criteria, leaving 404. The study was pre-registered.
Access to the system made people worse. Agreement with the system ran at 80.9 per cent with access against 58.4 per cent without, and accuracy at 63.9 per cent with against 74.2 per cent without, on a system that was right half the time. First-person hedging moved both in the right direction: agreement fell to 74.8 per cent and accuracy rose to 72.8 per cent, both significant. Impersonal hedging moved them in the same direction and did not reach significance. Breaking it down by whether the system was right, hedging "leads to some reduction in accuracy when the AI system is correct (92.2% to 89.2% for Uncertain1st) but a greater increase in accuracy when the AI system is incorrect (43.6% to 52.0%)".
Then the finding that explains why nobody ships this. Intention to use the system fell significantly under first-person hedging, from 3.25 to 2.91, while impersonal hedging left it at 3.36. So the version that helps the reader most is the version they least want to keep using, and the version they like is the one that does not work. A team optimising for engagement will select the impersonal hedge with no measurable benefit, and will be able to point at its uncertainty feature while doing it. The authors are careful about the limits of their own result, noting that their system "exhibited low accuracy and expressed uncertainty often, in a poorly calibrated manner", and that "regulators should avoid making blanket requirements on uncertainty expression, at least until more research has been done". That caution is worth reproducing rather than filtering out.
Silence is not neutral, and the confident register is trained in#
The strongest argument for requiring a marker rather than permitting one comes from Zhou, Hwang, Ren and Sap. They prompted nine models with 49 prompts over 284 questions, 125,244 queries in total, and found that "only 5% of the generated answers include any type of epistemic markers", so the overwhelming majority of what users see carries no signal at all. When models are pushed to express confidence they overshoot: an average 47 per cent error rate among responses expressed confidently, and only 53 per cent of generations expressing certainty are correct.
Their human study is the part to carry away. Participants shown a system's answers relied on hedged answers about 10 per cent of the time and confident ones about 90 per cent, which is the sensible response. But: "surprisingly, plain statements like 'The answer is (A)' or '(A)' are also relied on by users nearly 90% of the time. In other words, without communicating any epistemic markers, humans interpret this as a sign of model certainty." A system that says nothing about its uncertainty has not been neutral. It has asserted confidence, and been believed.
They also found the asymmetry recovers slowly. Participants exposed to an overconfident system and then to a calibrated one averaged 76 per cent during the miscalibrated rounds and 86 per cent afterwards, with the authors noting that "users' mental models were never fully corrected". An underconfident system produced the reverse: 66 per cent during, 98 per cent after. Being oversold takes longer to unlearn than being undersold.
And the mechanism sits in the training signal. Reward modelling scores plain statements at 4.03 on average, expressions of certainty at 0.82, and expressions of doubt at minus 1.86. Their reading: "there is not a human bias for strengtheners, rather there is a bias against weakeners." The confident register is what survives a preference model that punishes hedging harder than it rewards assurance, rather than a stylistic choice anyone made. That makes it a governance question about optimisation targets rather than a prompt-engineering question.
The law requires the disclosure and not the moment#
There is a common belief that European law now requires AI systems to communicate uncertainty. It does not. Article 13 of Regulation (EU) 2024/1689 requires providers of high-risk systems to supply deployers with instructions for use setting out "the level of accuracy, including its metrics, robustness and cybersecurity referred to in Article 15 against which the high-risk AI system has been tested and validated and which can be expected, and any known and foreseeable circumstances that may have an impact on that expected level of accuracy". Article 15(3) requires declared accuracy levels in those instructions. Article 50 requires that people be told they are interacting with an AI system, and that synthetic content be machine-readably marked.
All of that is documentary, aggregate and in advance, addressed to the deployer. The word uncertainty does not appear in Article 13, Article 15 or Article 50. Nothing in the Act requires a system to tell the person in front of it how sure it is about the specific thing it has just said. That is the entire subject of this page and it is unregulated, which makes it a design decision that organisations own rather than a compliance item they will be handed.
Four requirements worth writing into a specification#
These follow from the evidence above rather than from preference, and each is attributable. The assembly is this research's; the components belong to the people named.
- A fixed vocabulary with the range in the sentence. Not a glossary. Budescu and colleagues showed the glossary condition performed no better than nothing, and only in-text ranges helped. Borrow ICD 203's seven bands rather than inventing your own, because they have been in operational use for over a decade and inventing a scale is how you get a second incompatible scale.
- Likelihood and confidence stated separately. ICD 203's prohibition on combining them in one sentence is the single most portable rule in this territory. "Probably X, on a weak basis" and "possibly X, on a strong basis" are different messages, and a single confidence percentage cannot express either.
- First person, not impersonal. Kim and colleagues found the first-person hedge significantly reduced agreement with a wrong system and significantly raised reader accuracy, while the impersonal version reached neither. Accept the cost that came with it: intention to use fell. If your organisation is not prepared to pay that, it has decided against calibrated reliance and should say so out loud rather than shipping the version that tests well.
- No unmarked answers. Zhou and colleagues established that an unmarked statement is read as a confident one at close to the same rate as an explicitly confident one. Silence is therefore an assertion. Requiring a marker on every consequential output is the only way to make the absence of one mean anything.
One thing not on that list, deliberately. A confidence percentage on its own is weak. Zhang, Liao and Bellamy showed confidence scores significantly improved trust calibration and produced "no significant difference in AI-assisted accuracy across the prediction and confidence conditions", a result they report as rejecting their own hypothesis, and their local explanations did nothing at all. Their study is small, at nine participants per cell in the first experiment, so treat it as a caution rather than a settlement. But the direction is consistent with Steyvers on explanation length: making the interface say more about itself changes how the user feels rather than what they catch. This is the same ground as whether explaining an AI decision helps, and the answer there is the answer here.
Where this argument came from#
Rahim Hirji has been on the transparency question since before it was a product category. In AI Transparency in 2023 he argued that foundation-model opacity had become measurable and was failing on every axis, and that firms were setting the de facto rules ahead of anyone with the authority to write them. Three years later the calibration literature has caught up: the disclosure the Act requires is documentary, the per-output signal is unregulated, and the reward model decides.
The sharper connection is to Leave the Fingerprints In in 2026, whose argument is that the damage is in the reading rather than the writing, and that readers have grown a reflex for fluent, hollow prose. That is this page's problem stated as a literary one. Confidence in machine output is a property of the prose rather than of the knowledge. Hedging language therefore reads as weakness, and a well-written wrong answer outperforms a badly written right one. And in Rules Before Tools on 17 August 2025, the fourth of ten rules is titled "Make it legible", and its instruction is to attach plain-language explanations where money or people are affected and "set confidence thresholds for human review". The evidence above says the first half of that helps less than it appears to and the second half is where the value is.
What could not be confirmed for this page#
Several figures that circulate with this literature were checked and left out. Budescu, Broomell and Por's 2009 paper is the one everybody cites; its full text could not be opened, so its specific numbers do not appear and the 2012 nationally representative study by the same authors carries the argument instead. The 2014 multi-country replication is quoted from its abstract only, so no participant count is given. The expected calibration error values from Steyvers and colleagues sit inside a figure that did not survive text extraction, so only the AUC figures are used. The Zhang, Liao and Bellamy study reports F and p values without effect sizes or confidence intervals, and its cell sizes are very small.
Two structural gaps matter more. Nothing here tests whether calibrated uncertainty communication improves outcomes for domain experts, because every study above used lay participants on tasks they had no standing in. And the Zhou paper, which supplies the load-bearing finding about silence, reports no p values, confidence intervals or effect sizes for any of its human results, recruits US participants only, and describes its own view as narrow and US-centric. It earns its place on the size of its query count and the checkability of its reward-model measurement. The human half is 25 participants per setting and should be read as such.
Key sources
- Steyvers, M., Tejeda, H., Kumar, A., Belem, C., Karny, S., Hu, X., Mayer, L. W. and Smyth, P. (2025). What large language models know and what people think they know. Nature Machine Intelligence, 7, 221-231.
- Xiong, M., Hu, Z., Lu, X., Li, Y., Fu, J., He, J. and Hooi, B. (2024). Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs. ICLR 2024, arXiv:2306.13063.
- Kim, S. S. Y., Liao, Q. V., Vorvoreanu, M., Ballard, S. and Wortman Vaughan, J. (2024). "I'm Not Sure, But...": Examining the Impact of Large Language Models' Uncertainty Expression on User Reliance and Trust. FAccT 2024.
- Zhou, K., Hwang, J. D., Ren, X. and Sap, M. (2024). Relying on the Unreliable: The Impact of Language Models' Reluctance to Express Uncertainty. ACL 2024, 3623-3643. arXiv:2401.06730.
- Budescu, D. V., Por, H.-H. and Broomell, S. B. (2012). Effective communication of uncertainty in the IPCC reports. Climatic Change, 113, 181-200.
- Budescu, D. V., Por, H.-H., Broomell, S. B. and Smithson, M. (2014). The interpretation of IPCC probabilistic statements around the world. Nature Climate Change, 4(6), 508-512.
- Zhang, Y., Liao, Q. V. and Bellamy, R. K. E. (2020). Effect of confidence and explanation on accuracy and trust calibration in AI-assisted decision making. FAT* 2020.
- Office of the Director of National Intelligence. Intelligence Community Directive 203, Analytic Standards, signed 2 January 2015, technically amended 2022. The probability yardstick is in Tradecraft Standard 2(a).
- Kent, S. (1964). Words of Estimative Probability. Studies in Intelligence, 8(4). An essay rather than a study, and treated as one here.
- European Union (2024). Regulation (EU) 2024/1689, Article 13: Transparency and provision of information to deployers.
Related SuperSkills research#
On the confident register and where it comes from, why AI sounds so confident and what an AI hallucination is. On whether more interface helps, does explaining an AI decision help and automation bias. On what the reader is supposed to do with the signal, how to know when AI is wrong, when to override AI and algorithm aversion. On agents specifically, AI agents and human judgement, how humans and agents divide work across a process and who manages AI agents.
About this research#
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. The calibration gap and the discrimination gap are Steyvers and colleagues' terms. Words of estimative probability is Sherman Kent's. The probability yardstick belongs to Intelligence Community Directive 203. Calibration, expected calibration error and automation bias are established vocabulary. Nothing on this page is a SuperSkills coinage. Two author lists in wide circulation were corrected against the papers themselves during this build: the Nature Machine Intelligence paper has no author named Kerrigan or Askell, and the ICLR confidence-elicitation paper is Xiong, Hu, Lu, Li, Fu, He and Hooi. Articles 13, 15 and 50 were read at the European Commission's own AI Act Service Desk because EUR-Lex returned an empty document to every route tried; the Service Desk carries a notice that Article 50 has been amended by the Digital Omnibus and that its displayed text has not yet been updated, so the Article 50 wording above is the version of 13 June 2024.
Evidence review · SS-2026-170 · Graded against the published rubric
Hirji, R. (2026). How should an AI agent communicate uncertainty to a human?. The SuperSkills evidence base, SS-2026-170. https://thesuperskills.com/research/how-should-an-ai-agent-communicate-uncertainty. Last reviewed 4 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work