← Research
Research · Question

How do you build judgement in an organisation?

Four stages, each pairing a method that has been measured with an instrument that makes it visible.

Last reviewed: 26 September 2026 · Next review due: 26 September 2027

Most judgement training is awareness training, and awareness training has close to no evidence behind it. This is the alternative: four things that have been measured, sequenced, and joined to the instruments that produce a record. It is a programme rather than a course, and it runs on the work rather than beside it.

Question this page answersAll 996 questions this research covers

Four stages: make decisions visible, close the loop, train the two things that respond to training, and rebuild the ladder. Each stage pairs a method with measurement behind it against an instrument that leaves a record, and the order matters more than the content.

Share this line as a card

The answer, in one line

Parts of it, and less than most programmes claim.

Share as a card

What this programme is built against#

The judgement training on the market is mostly bias awareness, which has close to no behavioural evidence behind it. Accumulated experience does less than expected, because most professional work sits in what Hogarth called a wicked learning environment where feedback is delayed, noisy or censored by the decisions themselves. And the structured analytic techniques borrowed into business from intelligence work were reviewed in 2018 and found to have essentially no accuracy evidence.

What survives that clearing is set out at how judgement is trained: four methods with measurement behind them. This page sequences them and attaches each to the instrument that produces the record.

Stage one. Make the decision visible#

Weeks. Produces a record rather than an improvement, and everything after it depends on the record existing.

Nothing downstream works without this, because a decision nobody wrote down cannot be reviewed: hindsight rewrites the memory rather than suppressing it, so a team looking back at what they were thinking is examining a reconstruction organised around how it turned out.

What to expect. The first month is uncomfortable and produces nothing that looks like progress. Several teams will find they cannot name who made recent decisions. That finding is the output of stage one.

Stage two. Close the loop#

A quarter. The best-evidenced part of the programme and the cheapest.

Structured debriefs improved subsequent performance by around 20 to 25 per cent across 46 studies, with an average effect size near 0.67. The effect depends on conditions that are easy to drop: a real structure, participants reaching conclusions themselves rather than being told, a focus on process rather than outcome, and more than one person's account of what happened.

Stage three. Train the two things that respond to training#

Hours of instruction, months of scoring. Narrow, and the narrowness is the reason it works.

What to stop doing at this stage is as important: bias awareness sessions, and any programme whose output is a vocabulary rather than a practice.

Stage four. Rebuild the ladder#

Years. The slowest stage, the one that fails silently, and the one the other three exist to make possible.

Judgement has historically been built by doing small work badly under supervision, and that is the work being handed to machines first. The mechanism is at cognitive apprenticeship: expertise in cognitive work transfers when the reasoning is made explicit, and the machine's draft removes the occasion for anybody to make it explicit.

What to measure#

Four numbers, all cheap to collect and hard to fake, and none of them a score out of anything.

Two things not to measure. A score out of the seventy subskills, which produces a number with no instrument behind it. And anything that turns the five questions into a dashboard metric, since a measure gets optimised and those five do not survive it.

Where the Augmented Mindset sits in all of this#

The Augmented Mindset is what the four stages are developing. It is the one of the seven capabilities that shows up in a record rather than in an impression. A person who has it can produce the receipts, say what they changed in an output and why, name the terrain they were working on, and carry the way they work into an unfamiliar system. A person who has a configuration rather than a mindset can do none of those once the tool changes.

What this programme has not been shown to do#

It has not been run and evaluated as a programme. The four methods were measured separately, in different fields, on outcomes narrower than judgement at work: forecast accuracy, performance on a subsequent exercise, scores on a bias test, skills in a simulator. The step from those to better decisions in an organisation is an inference nobody has tested. The instruments from the book have no outcome evidence at all and are offered as ways of producing a record rather than as interventions. And the ordering is argued rather than trialled: it follows from the fact that review requires a record, which is reasoning, not a finding.

What can be said is that these are the only four methods with measurement behind them, that the alternative on the market has less, and that an organisation running this would at least know what it was doing and be able to show it.

Key sources

Evidence review · SS-2026-342 · Graded against the published rubric

Cite this page

Hirji, R. (2026). How do you build judgement in an organisation?. The SuperSkills evidence base, SS-2026-342. https://thesuperskills.com/research/how-do-you-build-judgement-in-an-organisation. Last reviewed 26 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

Can an organisation train judgement?

Parts of it, and less than most programmes claim. Four things have measurement behind them: structured debriefs, which improved subsequent performance by around a fifth across 46 studies; calibration training, worth 6 to 12 per cent in a randomised forecasting tournament for about an hour of instruction; a single well-designed debiasing exercise, with effects still present at twelve weeks; and simulation with graduated responsibility. Everything else on the market is awareness training, which has very little behind it.

What is the order?

Make decisions visible before trying to improve them, because a decision nobody recorded cannot be reviewed. Then close the loop with structured review. Then train the two narrow things that respond to training. Then rebuild the ladder, which is the slowest and the one that fails quietly if the first three are skipped.

How long does it take?

The first stage takes weeks and produces a record rather than an improvement. The second starts returning something within a quarter. The third is a matter of hours of instruction and months of scoring. The fourth is measured in years, because it is about how people become senior.

What should be measured?

Four things that are cheap to collect and hard to fake: the override rate, how often human review changes an output; the proportion of consequential decisions where a named person appears; the trend in recorded misses, since a record that stops finding them usually means somebody stopped looking; and calibration, scored in batches against forecasts that resolved.

In this hub

Definitions

The terms this field uses, defined against their primary sources.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory  ·  Boards  ·  Enquire