← Research
Research

What is judgement?

The thing this research says AI is coming for, defined so it can be argued with.

Last reviewed: 30 August 2026

A definition, three foundational accounts of where judgement comes from, the distinction between skill, capability and judgement, and what the evidence says happens to it under AI. Written because the argument on this site rests on a term it had never defined on its own terms.

Question this page answersAll 616 questions this research covers

Judgement is the capability to recognise what a situation is, and what it requires, before any option is weighed. It is built by accumulated exposure to cases rather than by instruction, much of it cannot be put into words. This is the part of expertise that shows up as knowing which question you are actually facing.

This research has argued for years that AI comes for judgement rather than for jobs, and has never defined the word on its own terms. That is a gap in the argument, and this page closes it so the claim can be disagreed with properly.

Experts do not compare options#

The most useful empirical account comes from Gary Klein, who studied fireground commanders, intensive-care nurses and military officers making decisions under time pressure. He expected to find people weighing alternatives. He found almost none.

Experts in real conditions recognise a situation as typical and generate a workable course of action directly, a process Klein named recognition-primed decision making. Sources of Power (1998). The expertise sits in the recognition, not in the comparison. By the time options are on the table, the hard part has happened.

Hubert and Stuart Dreyfus reached the same place developmentally. Their five-stage progression from novice to expert describes rule-following giving way to situational, intuitive judgement built from accumulated experience. Mind Over Machine (1986). The novice needs the rule because they cannot yet see the situation. The expert has stopped using the rule because they can.

Michael Polanyi explains why this is so hard to hand over. We can know more than we can tell. A large part of expert knowledge cannot be made explicit and is acquired through practice rather than instruction. The Tacit Dimension (1966).

Put the three together and you have a definition with consequences. Judgement is recognitional, developmental and largely tacit. It comes from having seen enough cases, it cannot be fully articulated, and it therefore cannot be restored by a training course once it has gone.

Skill, capability, judgement#

These three words are used interchangeably across most writing on this subject, including some of the writing on this site before now. They are not the same and the differences carry weight.

TermWhat it isHow it is acquiredHow it fails
SkillA trained, describable procedure with a criterion you can be tested againstInstruction plus practiceDecays measurably with disuse
JudgementRecognising what the situation is and what it requiresAccumulated exposure to cases, largely tacitNever forms, or degrades without anyone noticing
CapabilityWhat a person or organisation can actually do when the situation arrivesSkills, plus judgement, plus the conditions that let them be usedAny of the three going missing

The practical consequence is that an organisation can hold every relevant skill and lack the capability, because nobody present can tell which situation they are in. That is a different problem from a skills gap and it does not respond to the same remedy.

An organisation's capability is a further step again. It is what remains when particular individuals are not in the room: the precedents, the supervision, the escalation habits and the channels through which judgement passes from people who have it to people who do not. That last one is why tacit knowledge is the most exposed form of organisational knowledge. It moves through shared work, and shared work is what gets automated.

Judgement is not the decision#

The most common error in this field is treating judgement and decision-making as one thing. A decision is the moment of choosing between options that have already been framed. Judgement is what produced the framing: what kind of situation this is, which options deserve consideration, and what would count as a good outcome.

This distinction is what makes the oversight literature so uncomfortable. A person shown a well-argued recommendation can approve it, record that a human reviewed the decision, and never have exercised judgement at all, because the framing arrived with the recommendation. The estate treats this at human in the loop is not a safeguard and the difference between a good decision and a good outcome.

Why the evidence points at this specifically#

Two independent lines converge on judgement rather than on capability in general.

Vaccaro and colleagues meta-analysed 106 experiments and 370 effect sizes on human and AI combinations. The pairs performed worse on average than the better of either alone, Hedges' g = −0.23, and critically the losses were concentrated in decision-making while the gains were in content creation. Graded entry. The benchmark is an oracle-selected best performer that you rarely know in advance, which limits what the result prescribes. What it establishes is where the damage sits.

Arthur and colleagues, meta-analysing 189 data points on skill decay, found that cognitive, artificial and accuracy-based tasks decayed faster under non-use than physical, natural and speed-based ones. Graded entry. Casner's cockpit study found the same split by another route: manual control intact, and the situational tasks, tracking position and recognising failures, failing. Graded entry. More at how fast do skills decay.

One literature says human and AI pairings lose most where the work is deciding. Another says the faculties that decay fastest are the cognitive ones. Neither was designed to test this argument, and they meet on it.

The measured case is Budzyń's: nineteen endoscopists averaging 27.6 years of experience whose unassisted detection fell six percentage points within months of routine AI exposure. Graded entry. What degraded was recognition, which is Klein's definition of the thing.

Where the definition is contested#

Klein and Daniel Kahneman spent years on opposite sides of whether expert intuition should be trusted, then wrote up their disagreement together. Their conclusion in Conditions for Intuitive Expertise: A Failure to Disagree (American Psychologist, 2009) is that judging the likely quality of an intuitive judgement requires assessing the predictability of the environment in which it is made and the individual's opportunity to learn that environment's regularities. Firefighting supplies both. Long-horizon strategic forecasting supplies neither.

That matters here because it cuts against a comfortable reading of this whole research programme. If judgement is only dependable where the environment gives regular feedback, then in some domains a well-built system may make better calls than the expert, and protecting human judgement for its own sake would be sentimentality rather than strategy. The case for preserving judgement is therefore strongest exactly where recognition has been trained by real feedback, and weakest where the expert's confidence has never been tested against outcomes.

The Dreyfus model is also disputed. It is a phenomenological account rather than a measured one, and the five stages have never been validated as discrete. It is used here for the shape of the progression rather than as evidence for its steps.

What follows#

If judgement is recognitional, tacit and built by exposure, then three things follow that do not follow from treating it as a skill.

It cannot be trained back. A course transmits explicit knowledge, and the part that matters is the part that cannot be made explicit. Rebuilding judgement means rebuilding exposure to cases, which takes the time it originally took.

Its loss is invisible in output. The work still ships, because the framing arrived from somewhere. Only removing the tool reveals what is left, which is the argument in assessing capability rather than output.

The exposure has to be designed, because it used to be a by-product. Nobody arranged for juniors to see a thousand cases; it happened because the work passed through them. Remove the work and the exposure goes with it, which is the missing rungs and what accumulates is capability debt.

Key sources

The flagship argument is does AI reduce human judgement. On the parts of judgement that cannot be written down, tacit knowledge. On how it degrades, how fast do skills decay and deskilling. On why nobody notices, the illusion of competence. On the oversight that assumes it, human in the loop is not a safeguard and meaningful human oversight. On how it used to be built, the missing rungs and the missed reps. For the wider reading, the essential works.

About this research#

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. He set out the case for judgement as the scarce capability in The Next A.I. Power Class Won't Build the Models (Observer, 17 July 2026). The definition on this page is assembled from established work in decision research and the philosophy of knowledge; none of the three foundational accounts is a SuperSkills coinage.

How this research works  ·  Reviewed quarterly  ·  Found an error? Tell me and it is corrected on the page.

Cite this

Hirji, R. (2026). What is judgement? The SuperSkills Intelligence Company. Last reviewed 30 August 2026. thesuperskills.com/research/what-is-judgement

Questions answered on this page

What is judgement?

Judgement is the capability to recognise what a situation is, and what it requires, before any option is weighed. Gary Klein's studies of experts in real conditions found they rarely compare options at all: they recognise a situation as typical and generate a workable course of action directly. Hubert and Stuart Dreyfus described the same thing developmentally, as a five-stage progression in which rule-following gives way to situational, intuitive judgement built from accumulated experience. Michael Polanyi's point is why it resists training: we can know more than we can tell, so a large part of expert knowledge cannot be made explicit and is acquired through practice rather than instruction.

What is the difference between a skill and a capability?

A skill is a trained, describable procedure with a criterion you can be tested against. It can be taught by instruction and it decays measurably with disuse. A capability is what a person or organisation can actually do when a real situation arrives, which includes the relevant skills plus the judgement to recognise which apply, plus the conditions that let them be used. You can hold every relevant skill and still lack the capability, because nobody in the room can tell which situation they are in.

Is judgement the same as decision-making?

No, and conflating them is the most common error in this field. A decision is the moment of choosing between options that have already been framed. Judgement is what produces the framing: what kind of situation this is, which options are worth considering, and what would count as a good outcome. AI systems are increasingly good at the second half and are not making the first half. So a person can approve a well-argued recommendation to the wrong question and register it as having exercised oversight.

Why is judgement the thing at risk from AI?

Because of how it is built. Judgement comes from accumulated exposure to cases, much of it tacit, and it cannot be restored by a course. Two independent lines of evidence point at it specifically. In a meta-analysis of 106 experiments, human and AI combinations performed worse on average than the better of either alone, with the losses concentrated in decision-making and the gains in content creation. And in a meta-analysis of 189 data points on skill decay, cognitive and accuracy-based tasks decayed faster under non-use than physical and speed-based ones.

In this hub

Definitions

The terms this field uses, defined against their primary sources.

The work

Where the writing comes from.

These essays draw on research across more than 200 organisations in 30 countries. See the wider body of work, or bring it into your organisation.

All research →
Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory and coaching  ·  Enquire