← Research
Research

What is judgement?

The thing this research says AI is coming for, defined so it can be argued with.

Last reviewed: 30 August 2026

A definition, three foundational accounts of where judgement comes from, the distinction between skill, capability and judgement, and what the evidence says happens to it under AI. Written because the argument on this site rests on a term it had never defined on its own terms.

Judgement is the capability to recognise what a situation is, and what it requires, before any option is weighed. It is built by accumulated exposure to cases rather than by instruction, much of it cannot be put into words, and it is the part of expertise that shows up as knowing which question you are actually facing.

This research has argued for years that AI comes for judgement rather than for jobs, and has never defined the word on its own terms. That is a gap in the argument, and this page closes it so the claim can be disagreed with properly.

Experts do not compare options

The most useful empirical account comes from Gary Klein, who studied fireground commanders, intensive-care nurses and military officers making decisions under time pressure. He expected to find people weighing alternatives. He found almost none.

Experts in real conditions recognise a situation as typical and generate a workable course of action directly, a process Klein named recognition-primed decision making. Sources of Power (1998). The expertise sits in the recognition, not in the comparison. By the time options are on the table, the hard part has happened.

Hubert and Stuart Dreyfus reached the same place developmentally. Their five-stage progression from novice to expert describes rule-following giving way to situational, intuitive judgement built from accumulated experience. Mind Over Machine (1986). The novice needs the rule because they cannot yet see the situation. The expert has stopped using the rule because they can.

Michael Polanyi explains why this is so hard to hand over. We can know more than we can tell. A large part of expert knowledge cannot be made explicit and is acquired through practice rather than instruction. The Tacit Dimension (1966).

Put the three together and you have a definition with consequences. Judgement is recognitional, developmental and largely tacit. It comes from having seen enough cases, it cannot be fully articulated, and it therefore cannot be restored by a training course once it has gone.

Skill, capability, judgement

These three words are used interchangeably across most writing on this subject, including some of the writing on this site before now. They are not the same and the differences carry weight.

TermWhat it isHow it is acquiredHow it fails
SkillA trained, describable procedure with a criterion you can be tested againstInstruction plus practiceDecays measurably with disuse
JudgementRecognising what the situation is and what it requiresAccumulated exposure to cases, largely tacitNever forms, or degrades without anyone noticing
CapabilityWhat a person or organisation can actually do when the situation arrivesSkills, plus judgement, plus the conditions that let them be usedAny of the three going missing

The practical consequence is that an organisation can hold every relevant skill and lack the capability, because nobody present can tell which situation they are in. That is a different problem from a skills gap and it does not respond to the same remedy.

An organisation's capability is a further step again. It is what remains when particular individuals are not in the room: the precedents, the supervision, the escalation habits and the channels through which judgement passes from people who have it to people who do not. That last one is why tacit knowledge is the most exposed form of organisational knowledge. It moves through shared work, and shared work is what gets automated.

Judgement is not the decision

The most common error in this field is treating judgement and decision-making as one thing. A decision is the moment of choosing between options that have already been framed. Judgement is what produced the framing: what kind of situation this is, which options deserve consideration, and what would count as a good outcome.

This distinction is what makes the oversight literature so uncomfortable. A person shown a well-argued recommendation can approve it, record that a human reviewed the decision, and never have exercised judgement at all, because the framing arrived with the recommendation. The estate treats this at human in the loop is not a safeguard and the difference between a good decision and a good outcome.

Why the evidence points at this specifically

Two independent lines converge on judgement rather than on capability in general.

Vaccaro and colleagues meta-analysed 106 experiments and 370 effect sizes on human and AI combinations. The pairs performed worse on average than the better of either alone, Hedges' g = −0.23, and critically the losses were concentrated in decision-making while the gains were in content creation. Graded entry. The benchmark is an oracle-selected best performer that you rarely know in advance, which limits what the result prescribes. What it establishes is where the damage sits.

Arthur and colleagues, meta-analysing 189 data points on skill decay, found that cognitive, artificial and accuracy-based tasks decayed faster under non-use than physical, natural and speed-based ones. Graded entry. Casner's cockpit study found the same split by another route: manual control intact, and the situational tasks, tracking position and recognising failures, failing. Graded entry. More at how fast do skills decay.

One literature says human and AI pairings lose most where the work is deciding. Another says the faculties that decay fastest are the cognitive ones. Neither was designed to test this argument, and they meet on it.

The measured case is Budzyń's: nineteen endoscopists averaging 27.6 years of experience whose unassisted detection fell six percentage points within months of routine AI exposure. Graded entry. What degraded was recognition, which is Klein's definition of the thing.

Where the definition is contested

Klein and Daniel Kahneman spent years on opposite sides of whether expert intuition should be trusted, then wrote up their disagreement together. Their conclusion in Conditions for Intuitive Expertise: A Failure to Disagree (American Psychologist, 2009) is that judging the likely quality of an intuitive judgement requires assessing the predictability of the environment in which it is made and the individual's opportunity to learn that environment's regularities. Firefighting supplies both. Long-horizon strategic forecasting supplies neither.

That matters here because it cuts against a comfortable reading of this whole research programme. If judgement is only dependable where the environment gives regular feedback, then in some domains a well-built system may make better calls than the expert, and protecting human judgement for its own sake would be sentimentality rather than strategy. The case for preserving judgement is therefore strongest exactly where recognition has been trained by real feedback, and weakest where the expert's confidence has never been tested against outcomes.

The Dreyfus model is also disputed. It is a phenomenological account rather than a measured one, and the five stages have never been validated as discrete. It is used here for the shape of the progression rather than as evidence for its steps.

What follows

If judgement is recognitional, tacit and built by exposure, then three things follow that do not follow from treating it as a skill.

It cannot be trained back. A course transmits explicit knowledge, and the part that matters is the part that cannot be made explicit. Rebuilding judgement means rebuilding exposure to cases, which takes the time it originally took.

Its loss is invisible in output. The work still ships, because the framing arrived from somewhere. Only removing the tool reveals what is left, which is the argument in assessing capability rather than output.

The exposure has to be designed, because it used to be a by-product. Nobody arranged for juniors to see a thousand cases; it happened because the work passed through them. Remove the work and the exposure goes with it, which is the missing rungs and what accumulates is capability debt.

Key sources

Related SuperSkills research

The flagship argument is does AI reduce human judgement. On the parts of judgement that cannot be written down, tacit knowledge. On how it degrades, how fast do skills decay and deskilling. On why nobody notices, the illusion of competence. On the oversight that assumes it, human in the loop is not a safeguard and meaningful human oversight. On how it used to be built, the missing rungs and the missed reps. For the wider reading, the essential works.

About this research

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. He set out the case for judgement as the scarce capability in The Next A.I. Power Class Won't Build the Models (Observer, 17 July 2026). The definition on this page is assembled from established work in decision research and the philosophy of knowledge; none of the three foundational accounts is a SuperSkills coinage.

How this research works  ·  Reviewed quarterly  ·  Found an error? Tell me and it is corrected on the page.

Cite this

Hirji, R. (2026). What is judgement? The SuperSkills Intelligence Company. Last reviewed 30 August 2026. thesuperskills.com/research/what-is-judgement

The work

Where the writing comes from.

These essays draw on research across more than 200 organisations in 30 countries. See the wider body of work, or bring it into your organisation.

All research →
Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory and coaching  ·  Enquire