← Research
Research

Clinical versus actuarial judgement

The oldest quantitative result in this field, and the one most often used to settle arguments it does not settle.

Last reviewed: 26 September 2026 · Next review due: 26 September 2027

Meehl asked in 1954 whether a formula could beat an expert. Grove and colleagues answered it in 2000 across 136 studies. What the answer is, and what it does not license anybody to claim about AI.

Question this page answersAll 996 questions this research covers

Across seventy years of studies, simple mechanical rules equal or beat expert judgement more often than they lose to it, and the margin is far smaller than the way the finding is usually quoted. This is the oldest quantitative result in the judgement literature and the one most often used to settle arguments it does not settle.

The answer, in one line

On the same inputs, in settings with defined outcomes, mechanical rules equal or modestly beat expert judgement.

Share as a card

Definition#

Clinical versus actuarial judgement: the comparison between a human expert combining evidence in their head and a formula combining the same evidence mechanically. Clinical is the expert's integration of cues; actuarial, sometimes called mechanical or statistical, is any fixed rule applied consistently, including one built from very few variables.

Share this definition as a card

What the studies found#

Paul Meehl set the question in 1954 and expected the expert to win. Dawes, Faust and Meehl summarised the accumulated comparisons in Science in 1989. The most careful account is Grove and colleagues in 2000: a meta-analysis of 136 studies in which mechanical prediction was better in roughly half the comparisons, about equal in roughly half, and worse in a small minority, with an average advantage around ten per cent.

Read that shape carefully, because the headline usually loses it. The modal result is a tie. The advantage is real, consistent in direction and modest in size, and it holds across medicine, education, forecasting of violence, academic admission and job performance.

Why the formula wins when it wins#

Not because it knows more. It applies the same weights every time. The expert varies: across colleagues, across the week, across their own mood, which is the phenomenon a noise audit measures. Remove the variation and you gain accuracy without gaining any knowledge at all. That is the whole mechanism, and the reason a crude rule with three variables can match a specialist with twenty years of practice.

It also explains the boundary. Where the expert holds information the rule cannot see, and where the situation is unusual enough that the rule's weights no longer apply, the comparison stops favouring the machine. Meehl himself made the point with a broken leg: a formula predicting cinema attendance is useless in the week the subject is in plaster, and only a human knows about the plaster.

How it is misused in the AI conversation#

Three misreadings, in order of how often they arrive.

One, "algorithms beat experts". The evidence says mechanical combination usually equals and sometimes modestly exceeds expert combination, on the same inputs, in settings with clear outcomes. Half the comparisons are ties.

Two, treating a language model as an actuarial method. It is not one. An actuarial rule is transparent, fixed and reproducible; that is where its advantage comes from. A model that returns different output on the same input, and whose weights nobody can state, has the consistency property only partially and the transparency property not at all. The old finding does not transfer to it unexamined, and the literature has not yet examined it.

Three, using it against human oversight. The studies compare a human judging alone with a rule judging alone. They say almost nothing about a human reviewing a machine's output, which is the arrangement every organisation is actually building. What evidence exists on that combination is discouraging: see the meta-analysis at human and AI decision making.

What follows for an organisation#

Where a decision is repetitive, has a defined outcome and is currently made in somebody's head, the honest question is not whether to trust a model. It is whether the work would benefit from being made consistent at all, which is a cheaper and more testable question. A written rule, a checklist or a scoring sheet captures most of the available gain and can be inspected by anyone.

Where a decision is rare, consequential and unlike its predecessors, the finding does not apply and citing it is a rhetorical move rather than an argument. Most board decisions are in the second category, so this research treats governance of judgement separately from the automation of routine assessment.

Key sources

Essay · SS-2026-309

Cite this page

Hirji, R. (2026). Clinical versus actuarial judgement. The SuperSkills evidence base, SS-2026-309. https://thesuperskills.com/research/clinical-versus-actuarial-judgement. Last reviewed 26 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

Do algorithms make better decisions than experts?

On the same inputs, in settings with defined outcomes, mechanical rules equal or modestly beat expert judgement. Grove and colleagues, across 136 studies in 2000, found mechanical prediction better in roughly half the comparisons, about equal in roughly half, and worse in a small minority, with an average advantage of around ten per cent. The modal result is a tie.

Why does a simple rule match an expert?

Because it applies the same weights every time. The expert varies between colleagues, across the week and with their own mood. Removing that variation produces accuracy without adding any knowledge, and that is how a rule with three variables can match a specialist with twenty years of practice.

When does the human win?

Where the expert holds information the rule cannot see, and where the case is unusual enough that the rule's weights no longer apply. Meehl made the point with a broken leg: a formula predicting cinema attendance is useless in the week the subject is in plaster, and only a person knows about the plaster.

Does this research apply to AI?

Not directly. An actuarial rule is transparent, fixed and reproducible, and that is where its advantage comes from. A language model returns different output on the same input and its weights cannot be stated, so it has the consistency property only partially and the transparency property not at all.

Does it show that human oversight is unnecessary?

No. These studies compare a human judging alone with a rule judging alone. They say almost nothing about a human reviewing a machine's output, which is the arrangement organisations are actually building, and the evidence on that combination is discouraging.

In this hub

Definitions

The terms this field uses, defined against their primary sources.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory  ·  Boards  ·  Enquire