Ask most organisations what happens when a person disagrees with the system and you get a governance answer: there is an override, it sits with a named role, it is available at any time. Ask what happened the last time somebody used it and the room goes quiet. The interesting thing about disagreement between a person and a model is not how it gets resolved. It is that in most deployments it never surfaces at all, and the mechanism by which it fails to surface has now been measured.
The answer, in one line
In most deployments the disagreement never surfaces. A pre-registered experiment with 404 participants found that access to a system raised agreement from 58.4 to 80.9 per cent while lowering accuracy from 74.2 to 63.9 per cent.
The short answer#
Usually the disagreement disappears, resolved silently in the machine's favour, and the resolution is worse than either party alone would have managed. Three things drive it. The presence of an answer raises agreement and lowers accuracy at the same time. The system almost never signals when it is unsure. And the confidence information that would tell a person when to push back exists inside the model but does not reach them.
None of that is a reason to remove the human. It is a reason to notice that a conflict which never becomes visible has not been resolved, it has been absorbed.
What the presence of an answer does#
The cleanest measurement comes from a pre-registered experiment published at FAccT 2024: four conditions, eight yes-or-no medical questions, a final sample of 404 after pre-registered exclusions, with a fictional system whose answers were correct on exactly half the questions.
Access to the system raised agreement from 58.4 per cent to 80.9 per cent, and lowered accuracy from 74.2 per cent to 63.9 per cent. Both numbers moved together and in opposite directions: people agreed more and were right less.
The same study found something more useful than the problem. When the system expressed uncertainty in the first person, saying it was not sure, agreement fell to 74.8 per cent and accuracy rose to 72.8 per cent, both significant. Impersonal hedging moved the numbers the same way without reaching significance. So the wording of a machine's doubt changes whether a person keeps their own view.
The authors are careful about what follows, and so is this page: they say regulators should avoid blanket requirements to express uncertainty until more research is done, and note that their system had deliberately low accuracy. This is a finding about mechanism, not a policy recommendation.
How rarely the system says it is unsure#
If uncertainty expression is what preserves disagreement, the obvious question is how often it happens. A study across nine models, 49 prompts and 284 MMLU questions, 125,244 queries in total, found that only about 5 per cent of generated answers include any epistemic marker at all.
Among the answers expressed confidently, the error rate averages 47 per cent. Slightly better than half of everything stated with certainty is correct.
The human half of that study is weaker and should be read as suggestive: hedged answers were relied on around 10 per cent of the time and confident ones around 90, but plain unmarked statements were also relied on nearly 90 per cent of the time. The paper reports no total human sample size, no p values and no effect sizes for the human results, with 25 participants per setting. It is quoted here for the model-side numbers, which are large and well specified, and not for the behavioural claim.
The signal exists and does not arrive#
This is the finding that makes the whole thing tractable. Two behavioural experiments published in Nature Machine Intelligence, 301 participants, compared how well model confidence separates right answers from wrong ones against how well a person can do the same after reading the model's explanation.
Model confidence discriminates correct from incorrect at an AUC of 0.751 for GPT-3.5, 0.746 for PaLM2 and 0.781 for GPT-4o. Participants reading the default explanations reached 0.589, 0.602 and 0.592, which the authors describe as only slightly better than random guessing. They name the two shortfalls the calibration gap and the discrimination gap, and attribute the human side primarily to overconfidence.
Put plainly: the system holds usable information about when it is likely to be wrong, and the person reading its output cannot recover that information from what they are shown. A reader deciding whether to disagree is working from something close to noise.
The authors note that closing the gap has not been shown to improve task accuracy; their improved-explanation result is a post-hoc simulation rather than a live deployment, and their participants had no domain expertise.
Aviation had this problem and named it#
The pattern of a junior party holding a correct view that never reaches the decision is not new, and one industry took it seriously enough to build a discipline around it. Crew Resource Management began with a 1979 NASA workshop prompted by an NTSB finding that a captain had failed to accept input from junior crew.
What CRM did was make disagreement a procedure rather than a personality trait: who is expected to speak, in what words, at what point, and what the senior party must do in response. Line audits show it produces the intended behavioural change, though measured attitudes decay over time even with recurrent training. The authors are explicit that it has not been shown to reduce accidents, because accidents are too rare to serve as a validation criterion.
What transfers is not the checklist but the recognition underneath it: a disagreement that depends on somebody feeling brave enough to voice it is one most organisations will never hear.
Making the disagreement visible#
- Record the disagreement, not just the decision. If the only thing logged is the outcome, an organisation cannot tell the difference between agreement and absorption. A field for "did the reviewer differ, and what happened" costs nothing and is the only way this becomes visible.
- Ask when the override was last used. Not whether one exists. A control nobody has exercised is untested, and the aviation literature suggests the reason is usually social rather than technical.
- Treat the absence of hedging as uninformative. Roughly 95 per cent of answers carry no epistemic marker, so a confident tone tells you nothing about whether this is one of the 47 per cent.
- Do not rely on the explanation to calibrate the reader. The evidence says people reading default explanations are close to guessing about correctness. If a decision needs calibration, it has to come from something other than the model's own account of itself.
Whether a system should say it is unsure#
Whether expressed uncertainty should be required is open, and the researchers closest to the evidence say so. There is a real risk on the other side: a system that hedges constantly trains people to ignore the hedge, and one of these studies found intention to use fell significantly when the system expressed doubt. A tool nobody wants to use has solved the disagreement problem in the least useful way.
The studies here also sit mostly in short factual tasks with participants who were not domain experts. What happens when a specialist disagrees with a model in their own field, repeatedly, over months, is the question that matters most in professional work and the one nobody has run.
Evidence review · SS-2026-205 · Graded against the published rubric
Hirji, R. (2026). What happens when AI and human judgement conflict?. The SuperSkills evidence base, SS-2026-205. https://thesuperskills.com/research/ai-and-human-disagreement. Last reviewed 9 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work