Across seventy years of studies, simple mechanical rules equal or beat expert judgement more often than they lose to it, and the margin is far smaller than the way the finding is usually quoted. This is the oldest quantitative result in the judgement literature and the one most often used to settle arguments it does not settle.
The answer, in one line
On the same inputs, in settings with defined outcomes, mechanical rules equal or modestly beat expert judgement.
Definition#
Clinical versus actuarial judgement: the comparison between a human expert combining evidence in their head and a formula combining the same evidence mechanically. Clinical is the expert's integration of cues; actuarial, sometimes called mechanical or statistical, is any fixed rule applied consistently, including one built from very few variables.
What the studies found#
Paul Meehl set the question in 1954 and expected the expert to win. Dawes, Faust and Meehl summarised the accumulated comparisons in Science in 1989. The most careful account is Grove and colleagues in 2000: a meta-analysis of 136 studies in which mechanical prediction was better in roughly half the comparisons, about equal in roughly half, and worse in a small minority, with an average advantage around ten per cent.
Read that shape carefully, because the headline usually loses it. The modal result is a tie. The advantage is real, consistent in direction and modest in size, and it holds across medicine, education, forecasting of violence, academic admission and job performance.
Why the formula wins when it wins#
Not because it knows more. It applies the same weights every time. The expert varies: across colleagues, across the week, across their own mood, which is the phenomenon a noise audit measures. Remove the variation and you gain accuracy without gaining any knowledge at all. That is the whole mechanism, and the reason a crude rule with three variables can match a specialist with twenty years of practice.
It also explains the boundary. Where the expert holds information the rule cannot see, and where the situation is unusual enough that the rule's weights no longer apply, the comparison stops favouring the machine. Meehl himself made the point with a broken leg: a formula predicting cinema attendance is useless in the week the subject is in plaster, and only a human knows about the plaster.
How it is misused in the AI conversation#
Three misreadings, in order of how often they arrive.
One, "algorithms beat experts". The evidence says mechanical combination usually equals and sometimes modestly exceeds expert combination, on the same inputs, in settings with clear outcomes. Half the comparisons are ties.
Two, treating a language model as an actuarial method. It is not one. An actuarial rule is transparent, fixed and reproducible; that is where its advantage comes from. A model that returns different output on the same input, and whose weights nobody can state, has the consistency property only partially and the transparency property not at all. The old finding does not transfer to it unexamined, and the literature has not yet examined it.
Three, using it against human oversight. The studies compare a human judging alone with a rule judging alone. They say almost nothing about a human reviewing a machine's output, which is the arrangement every organisation is actually building. What evidence exists on that combination is discouraging: see the meta-analysis at human and AI decision making.
What follows for an organisation#
Where a decision is repetitive, has a defined outcome and is currently made in somebody's head, the honest question is not whether to trust a model. It is whether the work would benefit from being made consistent at all, which is a cheaper and more testable question. A written rule, a checklist or a scoring sheet captures most of the available gain and can be inspected by anyone.
Where a decision is rare, consequential and unlike its predecessors, the finding does not apply and citing it is a rhetorical move rather than an argument. Most board decisions are in the second category, so this research treats governance of judgement separately from the automation of routine assessment.
Key sources
- Grove, W. M., Zald, D. H., Lebow, B. S., Snitz, B. E. and Nelson, C. (2000). Clinical versus mechanical prediction: a meta-analysis. Psychological Assessment, 12(1), 19-30.
- Dawes, R. M., Faust, D. and Meehl, P. E. (1989). Clinical versus actuarial judgment. Science, 243(4899), 1668-1674.
- Meehl, P. E. (1954). Clinical versus Statistical Prediction. University of Minnesota Press.
- Vaccaro, M., Almaatouq, A. and Malone, T. (2024). When combinations of humans and AI are useful. Nature Human Behaviour, 8.
Essay · SS-2026-309
Hirji, R. (2026). Clinical versus actuarial judgement. The SuperSkills evidence base, SS-2026-309. https://thesuperskills.com/research/clinical-versus-actuarial-judgement. Last reviewed 26 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work