Two well-run sets of experiments reached opposite conclusions about whether people take algorithmic advice. One found they take too little of it. The other found they take too much. Both are correct, and the thing that separates them is the answer to this question: expertise. The person who gains most from AI advice is the one who knows least, and the person being asked to defer is usually the one who knows most. That inversion is the whole difficulty, and almost no adoption programme is designed around it.
The answer, in one line
It depends who is being assisted. For a novice, algorithmic advice frequently improves the decision and the measured gains are large. For a domain expert the average gain is close to nothing and the variance between individuals is enormous.
The short answer#
It depends who "we" is, and the dependence runs the opposite way to the way it is usually assumed. For a novice in a domain, algorithmic advice frequently improves the decision and the measured gains are large. For a domain expert, the average gain is close to nothing, the variance between individuals is enormous, and there is currently no reliable way to tell in advance which experts will be helped and which will be harmed.
So "trust AI over human experts" is not a policy. The defensible policy is narrower: work out what a person could have detected unaided, and let that decide how much weight the advice carries.
The two findings that look contradictory#
Six experiments on estimates and forecasts, published in Organizational Behavior and Human Decision Processes, found what the authors called algorithm appreciation: people often weight algorithmic advice more heavily than advice from another person. Domain experts were the notable exception.
Five experiments published four years earlier found algorithm aversion: after seeing an algorithm err, people abandon it even when it demonstrably outperforms them, and forgive the identical mistake in a human.
The two results sit together once you notice they are describing different people at different moments. Before an error is visible, and in a domain where the person has no standing of their own, the algorithm is preferred. After a visible error, or where the person has expertise to defend, it is discounted. Neither study tells you which tendency will dominate in a particular workplace, and the authors of the first say so directly.
What happens when the expert is the one being assisted#
The best-designed study on this point is a randomised experiment across 140 radiologists published in Nature Medicine: fifteen chest X-ray diagnostic tasks, roughly 5,190 observations, with empirical-Bayes shrinkage applied to separate genuine individual differences from noise.
The finding is not that AI assistance helped or that it did not. It is that the effect diverged sharply between individual radiologists, from strongly positive to strongly negative, and that experience, subspecialty and prior familiarity with AI all failed to predict who would benefit. Lower performers did not consistently gain either, which removes the intuitive fallback of aiming assistance at the weakest readers.
Read that carefully, because it is stronger than a null result. A null result would mean assistance does nothing. This means assistance does something large, in both directions, to people you cannot tell apart in advance. Any policy of the form "clinicians should defer to the model" is therefore making a bet on an unobserved individual characteristic.
The authors are explicit about what it does not show: not that AI assistance is bad on average, and not that the pattern holds outside diagnostic imaging.
The order in which they see it changes the answer#
There is a mechanical finding underneath the psychological one. A between-subjects study of 19 veterinary radiologists compared showing the AI output alongside the image against having the clinician commit to a provisional diagnosis first. Final diagnoses matched the AI 91 per cent of the time when the AI was seen first, against 89 per cent when the clinician committed first. Where the AI had flagged a finding, agreement ran at 71 per cent against 65 per cent.
Those are small numbers on a small sample in one speciality, and the study is a working paper rather than a peer-reviewed result, so it earns a note rather than an argument. What it points at is worth the note: the anchoring produced only marginal diagnostic gains, because agreement rose on the erroneous advice as well as the correct advice. Sequencing is a design decision that most deployments make by accident.
Why the gains land on the novice#
The economics runs the same way. A staggered rollout across 5,172 customer-support agents, three million chats, published in the Quarterly Journal of Economics, found resolutions per hour rose 15 per cent on average. The average conceals the result: less skilled and less experienced workers gained 30 per cent, rising to 36 per cent in the lowest skill quintile, while the most skilled saw no significant productivity gain at all.
The same shape appears well outside white-collar work. Driver-level data from a Japanese taxi fleet through the rollout of an AI demand-prediction system found the gains accrued almost entirely to low-skilled drivers, narrowing the gap between best and worst by 14 per cent. That entry sits in this evidence base deliberately as a disconfirming case: anyone arguing that AI levels up office workers while degrading frontline ones has to explain a result running the other way.
Put the three together and the picture is consistent. AI advice substitutes for expertise the person does not have. Where the expertise is already there, it has much less to add, and what it adds is unpredictable at the individual level.
The cost that arrives later#
There is a second reason not to convert this into a rule about deference. A study of endoscopists found that adenoma detection rate in unassisted colonoscopy fell from 28.4 per cent before AI exposure to 22.4 per cent after, a drop of six percentage points. The assistance improved the assisted reading and degraded the unassisted one.
That is the shape of the trade. A policy of deferring to the model is not only a decision about today's accuracy. It is a decision about what the expert will be able to do on the day the model is unavailable, wrong, or facing a case outside its distribution.
Who is being assisted, and in what order#
- Ask who is being assisted, not whether the tool is good. The same system deployed to novices and to experts is two different interventions with two different expected returns.
- Decide the sequence deliberately. Commit-first costs time and preserves an independent read. Advice-first is faster and anchors. Both are legitimate; picking one by accident is not.
- Measure the unassisted case. If nobody tests what people can do without the system, the only capability you are tracking is the one that includes it.
- Stop looking for the type of expert who benefits. The best available study looked and could not find one. Design for the variance instead of trying to predict it.
Three domains with checkable answers, and the ones without#
All of the strongest evidence here comes from diagnostic imaging, customer support and taxi dispatch. Those are domains with reasonably fast feedback and a checkable ground truth. That is what makes them measurable, and the same property makes them unrepresentative of law, strategy or management, where nobody has run the study.
Nor does any of this say the expert is right. It says the expert is the person whose independent read is expensive to reconstruct once it has gone, and that treating deference as a default spends that read without recording the cost.
Evidence review · SS-2026-204 · Graded against the published rubric
Hirji, R. (2026). Should we trust AI over human experts?. The SuperSkills evidence base, SS-2026-204. https://thesuperskills.com/research/ai-and-expert-judgement. Last reviewed 9 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work