← Research
Research

The models of judgement

Twenty frameworks, five methods, and one thing none of them does.

Last reviewed: 26 September 2026 · Next review due: 26 September 2027

Noise audits, the decision quality chain, clinical versus actuarial prediction, recognition-primed decisions, levels of automation, the lens model, prediction and judgement, and the business frames from London Business School to the World Economic Forum. Each with what it establishes and what it does not.

Question this page answersAll 996 questions this research covers

There are about twenty named models of human judgement in circulation, they disagree with each other, and only five of them give an organisation anything it can actually run. This page lists them with what each one says, what it does not establish, and whether it is a method or a vocabulary. It exists because the alternative is a conversation in which everybody cites a framework and nobody says what it proves.

Share this line as a card

The answer, in one line

Five give you a method: the noise audit, the Mediating Assessments Protocol, the decision quality chain, the four-stage types and levels of automation model, and situation awareness with SAGAT.

Share as a card

How this page grades them#

Three questions per model. Who established it and when. What it says, in one sentence, without the gloss it has acquired since. And whether it hands you a protocol, a measurable or a set of steps, as against a description you can nod at. The third question is the one that sorts the list, and most of the list fails it.

The models that give you a method#

The noise audit. Kahneman, Sibony and Sunstein, 2021. Several professionals judge the same real cases independently and you measure the scatter. Gives you a number before anybody argues. Does not establish that consistency is accuracy, and a quiet biased process is worse than a noisy one. The full account.

The Mediating Assessments Protocol. Kahneman, Lovallo and Sibony, 2019, in MIT Sloan Management Review. Break a decision into independent assessments, score each separately, and hold the overall judgement until the end. Borrowed from structured interviewing, where the evidence is strong; never trialled on strategic decisions, where it is recommended.

The decision quality chain. Spetzler, Winter and Meyer, 2016, on Howard's decision analysis. Six links, and the quality of the decision is the weakest of them. A practitioner framework with no controlled trial behind it, and the weakest-link property is what earns its place. The six links.

Types and levels of automation. Parasuraman, Sheridan and Wickens, 2000. Four stages of work, each automated to a chosen degree, evaluated against consequences for the human. The closest thing to an allocation instrument that exists. Contains no rule for choosing the level. How the four stages work.

Situation awareness and SAGAT. Endsley, 1995. Perception, comprehension, projection, with a measurement procedure that stops the task and asks what the operator knows. The construct is contested and the measure needs a bounded task, which knowledge work rarely has.

The models that give you a vocabulary#

Clinical versus actuarial judgement. Meehl 1954, Dawes, Faust and Meehl 1989, Grove and colleagues 2000. Mechanical combination equals or modestly beats expert combination; the modal result is a tie. Says nothing about a human reviewing a machine, which is what everyone is building. What it does and does not show.

The recognition-primed decision model. Klein, 1993. Experts recognise a situation and simulate the first workable option rather than comparing alternatives. Valid in high-validity, fast-feedback environments, which Kahneman and Klein jointly specified in 2009; most management decisions are not those environments.

The lens model. Brunswik, 1952, formalised later. Accuracy decomposes into how predictable the world is, how consistent the judge is and how well their cue weights match. An accounting identity rather than a theory, and it needs outcome data most organisations do not keep.

The cognitive continuum. Hammond, 1996. Cognition runs from intuition to analysis and the task should induce the matching mode. No validated way of saying where a given task sits.

Skill, rule and knowledge. Rasmussen, 1983. Three levels of performance with distinct error types. A taxonomy, widely used in safety engineering, with no predictive rates attached.

Ironies of automation. Bainbridge, 1983. Automating the easy parts leaves the human the hard residue with less practice at it. Four pages, no data, and the most quoted argument in this field. It is an argument, and this estate treats it as one.

Prediction and judgement. Agrawal, Gans and Goldfarb, 2018. Cheap prediction raises the value of judgement, defined as specifying the payoff. An economic model rather than a measurement. It is routinely quoted as a finding.

The findings that complicate all of it#

Human and AI combinations usually underperform the better of the two. Vaccaro, Almaatouq and Malone, 2024, across 106 experiments, in Nature Human Behaviour. Synergy appeared on creation tasks and not on decision tasks. This is the single most load-bearing result in the territory and it is increasingly cited as the opposite of what it says.

Explanations increase acceptance without improving accuracy. Bansal and colleagues, CHI 2021. The intuitive fix for oversight makes the measured outcome worse. See appropriate reliance.

Aversion and appreciation both exist. Dietvorst 2015 found people abandon an algorithm after seeing it err; Logg 2019 found lay people weight algorithmic advice more heavily than human advice. Both predate large language models and neither settles the question for them.

The frames from business writing, named honestly#

These are not research findings and they are the ones a leadership team is most likely to have met. Andrew Likierman's six elements of judgment, from London Business School: knowledge, context, trust, feelings, choice, delivery. The World Economic Forum's "judgement work". MIT Sloan Management Review's intelligent choice architectures. Deloitte's use of the one-way and two-way door distinction. They are useful vocabulary, they are largely untested, and they share an assumption this research disputes: that judgement is a personal capacity leaders should sharpen.

What none of them do#

Not one of these models measures judgement in open-ended work with a language model in the loop. There is no validated instrument, no agreed unit and no longitudinal study of an organisation losing the capacity. Everything current is self-report, benchmark accuracy or reliance rate, and each of the three measures something adjacent.

Second gap, and the one this estate is built on: every framework above treats judgement as a property of a person. None of them treats it as a property of an organisation, which is where it is actually being lost, by a thousand reasonable allocation decisions nobody made on purpose. That argument is at human judgement in the age of AI and the mechanism is at drift versus design.

The definition this research uses is at what is judgement. The evidence on whether AI weakens it is at AI and human judgement. The measured record on combinations is at human and AI decision making. The whole graded base is at the evidence base.

Essay · SS-2026-312

Cite this page

Hirji, R. (2026). The models of judgement. The SuperSkills evidence base, SS-2026-312. https://thesuperskills.com/research/models-of-judgement. Last reviewed 26 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

What are the main models of human judgement?

Five give you a method: the noise audit, the Mediating Assessments Protocol, the decision quality chain, the four-stage types and levels of automation model, and situation awareness with SAGAT. The rest give you vocabulary: clinical versus actuarial prediction, recognition-primed decisions, the lens model, the cognitive continuum, skill-rule-knowledge, the ironies of automation, and the prediction-versus-judgement distinction from economics.

Which model of judgement is most useful to an organisation?

The one that produces a number it did not have. A noise audit measures disagreement between people doing the same job in an afternoon. The decision quality chain finds which of six links is failing. The four-stage automation model turns how much AI into four separate questions with four separate answers.

Do human and AI combinations make better decisions?

On average, no. Vaccaro, Almaatouq and Malone examined 106 experiments in Nature Human Behaviour in 2024 and found human and AI combinations performing worse than the better of the two alone, with synergy appearing on creation tasks and not on decision tasks. It is the most load-bearing recent result in this field and it is frequently cited as the opposite of what it says.

What do the business frameworks add?

Vocabulary a leadership team is likely to have met: Andrew Likierman's six elements of judgment at London Business School, the World Economic Forum's judgement work, MIT Sloan Management Review's intelligent choice architectures, and Deloitte's use of one-way and two-way doors. They are useful and largely untested, and they share the assumption that judgement is a personal capacity leaders should sharpen.

What do none of these models do?

None measures judgement in open-ended work with a language model in the loop: there is no validated instrument, no agreed unit and no longitudinal study of an organisation losing the capacity. And every one of them treats judgement as a property of a person rather than of an organisation, which is where it is actually being lost.

In this hub

Definitions

The terms this field uses, defined against their primary sources.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory  ·  Boards  ·  Enquire