There are about twenty named models of human judgement in circulation, they disagree with each other, and only five of them give an organisation anything it can actually run. This page lists them with what each one says, what it does not establish, and whether it is a method or a vocabulary. It exists because the alternative is a conversation in which everybody cites a framework and nobody says what it proves.
The answer, in one line
Five give you a method: the noise audit, the Mediating Assessments Protocol, the decision quality chain, the four-stage types and levels of automation model, and situation awareness with SAGAT.
How this page grades them#
Three questions per model. Who established it and when. What it says, in one sentence, without the gloss it has acquired since. And whether it hands you a protocol, a measurable or a set of steps, as against a description you can nod at. The third question is the one that sorts the list, and most of the list fails it.
The models that give you a method#
The noise audit. Kahneman, Sibony and Sunstein, 2021. Several professionals judge the same real cases independently and you measure the scatter. Gives you a number before anybody argues. Does not establish that consistency is accuracy, and a quiet biased process is worse than a noisy one. The full account.
The Mediating Assessments Protocol. Kahneman, Lovallo and Sibony, 2019, in MIT Sloan Management Review. Break a decision into independent assessments, score each separately, and hold the overall judgement until the end. Borrowed from structured interviewing, where the evidence is strong; never trialled on strategic decisions, where it is recommended.
The decision quality chain. Spetzler, Winter and Meyer, 2016, on Howard's decision analysis. Six links, and the quality of the decision is the weakest of them. A practitioner framework with no controlled trial behind it, and the weakest-link property is what earns its place. The six links.
Types and levels of automation. Parasuraman, Sheridan and Wickens, 2000. Four stages of work, each automated to a chosen degree, evaluated against consequences for the human. The closest thing to an allocation instrument that exists. Contains no rule for choosing the level. How the four stages work.
Situation awareness and SAGAT. Endsley, 1995. Perception, comprehension, projection, with a measurement procedure that stops the task and asks what the operator knows. The construct is contested and the measure needs a bounded task, which knowledge work rarely has.
The models that give you a vocabulary#
Clinical versus actuarial judgement. Meehl 1954, Dawes, Faust and Meehl 1989, Grove and colleagues 2000. Mechanical combination equals or modestly beats expert combination; the modal result is a tie. Says nothing about a human reviewing a machine, which is what everyone is building. What it does and does not show.
The recognition-primed decision model. Klein, 1993. Experts recognise a situation and simulate the first workable option rather than comparing alternatives. Valid in high-validity, fast-feedback environments, which Kahneman and Klein jointly specified in 2009; most management decisions are not those environments.
The lens model. Brunswik, 1952, formalised later. Accuracy decomposes into how predictable the world is, how consistent the judge is and how well their cue weights match. An accounting identity rather than a theory, and it needs outcome data most organisations do not keep.
The cognitive continuum. Hammond, 1996. Cognition runs from intuition to analysis and the task should induce the matching mode. No validated way of saying where a given task sits.
Skill, rule and knowledge. Rasmussen, 1983. Three levels of performance with distinct error types. A taxonomy, widely used in safety engineering, with no predictive rates attached.
Ironies of automation. Bainbridge, 1983. Automating the easy parts leaves the human the hard residue with less practice at it. Four pages, no data, and the most quoted argument in this field. It is an argument, and this estate treats it as one.
Prediction and judgement. Agrawal, Gans and Goldfarb, 2018. Cheap prediction raises the value of judgement, defined as specifying the payoff. An economic model rather than a measurement. It is routinely quoted as a finding.
The findings that complicate all of it#
Human and AI combinations usually underperform the better of the two. Vaccaro, Almaatouq and Malone, 2024, across 106 experiments, in Nature Human Behaviour. Synergy appeared on creation tasks and not on decision tasks. This is the single most load-bearing result in the territory and it is increasingly cited as the opposite of what it says.
Explanations increase acceptance without improving accuracy. Bansal and colleagues, CHI 2021. The intuitive fix for oversight makes the measured outcome worse. See appropriate reliance.
Aversion and appreciation both exist. Dietvorst 2015 found people abandon an algorithm after seeing it err; Logg 2019 found lay people weight algorithmic advice more heavily than human advice. Both predate large language models and neither settles the question for them.
The frames from business writing, named honestly#
These are not research findings and they are the ones a leadership team is most likely to have met. Andrew Likierman's six elements of judgment, from London Business School: knowledge, context, trust, feelings, choice, delivery. The World Economic Forum's "judgement work". MIT Sloan Management Review's intelligent choice architectures. Deloitte's use of the one-way and two-way door distinction. They are useful vocabulary, they are largely untested, and they share an assumption this research disputes: that judgement is a personal capacity leaders should sharpen.
What none of them do#
Not one of these models measures judgement in open-ended work with a language model in the loop. There is no validated instrument, no agreed unit and no longitudinal study of an organisation losing the capacity. Everything current is self-report, benchmark accuracy or reliance rate, and each of the three measures something adjacent.
Second gap, and the one this estate is built on: every framework above treats judgement as a property of a person. None of them treats it as a property of an organisation, which is where it is actually being lost, by a thousand reasonable allocation decisions nobody made on purpose. That argument is at human judgement in the age of AI and the mechanism is at drift versus design.
Related SuperSkills research#
The definition this research uses is at what is judgement. The evidence on whether AI weakens it is at AI and human judgement. The measured record on combinations is at human and AI decision making. The whole graded base is at the evidence base.
Essay · SS-2026-312
Hirji, R. (2026). The models of judgement. The SuperSkills evidence base, SS-2026-312. https://thesuperskills.com/research/models-of-judgement. Last reviewed 26 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work