Four stages: make decisions visible, close the loop, train the two things that respond to training, and rebuild the ladder. Each stage pairs a method with measurement behind it against an instrument that leaves a record, and the order matters more than the content.
What this programme is built against#
The judgement training on the market is mostly bias awareness, which has close to no behavioural evidence behind it. Accumulated experience does less than expected, because most professional work sits in what Hogarth called a wicked learning environment where feedback is delayed, noisy or censored by the decisions themselves. And the structured analytic techniques borrowed into business from intelligence work were reviewed in 2018 and found to have essentially no accuracy evidence.
What survives that clearing is set out at how judgement is trained: four methods with measurement behind them. This page sequences them and attaches each to the instrument that produces the record.
Stage one. Make the decision visible#
Weeks. Produces a record rather than an improvement, and everything after it depends on the record existing.
Nothing downstream works without this, because a decision nobody wrote down cannot be reviewed: hindsight rewrites the memory rather than suppressing it, so a team looking back at what they were thinking is examining a reconstruction organised around how it turned out.
- Map the terrain first. The three terrains, written before the tools go in: the patterns the system is trusted on, the calls that are kept, and the decisions where an override always runs.
- Keep the receipts. Hand, Head, Hours, Heart on consequential decisions. Four short lines, written as the work happens.
- Add a decision journal for the class of decisions that would be expensive to get wrong, carrying a confidence figure and a falsifier. Five honest entries a month beat fifty dutiful ones.
- Write the custody list. Custody is decided before the tools arrive, because afterwards the answer bends towards whatever is convenient.
What to expect. The first month is uncomfortable and produces nothing that looks like progress. Several teams will find they cannot name who made recent decisions. That finding is the output of stage one.
Stage two. Close the loop#
A quarter. The best-evidenced part of the programme and the cheapest.
Structured debriefs improved subsequent performance by around 20 to 25 per cent across 46 studies, with an average effect size near 0.67. The effect depends on conditions that are easy to drop: a real structure, participants reaching conclusions themselves rather than being told, a focus on process rather than outcome, and more than one person's account of what happened.
- Run the review properly. The after action review, held soon enough that memory has not reorganised, in a room where naming a mistake costs nothing.
- Add the sixty-second version everywhere else. The pause minute, callable by anybody rather than run by the chair.
- Put the loops on a cadence. The Five Loops supply the closure signals, and each of them requires a human to appear by name, which is what makes the rhythm resistant to usage theatre.
- Review the process, not the result. Being told the answer was wrong teaches very little. Being shown how you weighted the available information is what improves judgement.
Stage three. Train the two things that respond to training#
Hours of instruction, months of scoring. Narrow, and the narrowness is the reason it works.
- Calibration. About an hour of probabilistic reasoning training improved accuracy by roughly 6 to 12 per cent against control in a randomised tournament, replicated across years. Calibration training works because the forecasts resolve, so the habit to install is stating confidence as a number and scoring it in batches.
- One structured debiasing exercise. A single interactive session with immediate feedback on the person's own errors produced large reductions still present at twelve weeks, and outperformed a video on the same content. Practice with feedback, not information. See do debiasing programmes work.
- Two forcing functions from the book. The Principle Test before decisions taken under pressure, and the premortem before commitments. Neither has outcome evidence; both are cheap and both license doubts the room otherwise suppresses.
- At the task level. The use-or-keep test, for the individual deciding what to hand over on a specific piece of work.
What to stop doing at this stage is as important: bias awareness sessions, and any programme whose output is a vocabulary rather than a practice.
Stage four. Rebuild the ladder#
Years. The slowest stage, the one that fails silently, and the one the other three exist to make possible.
Judgement has historically been built by doing small work badly under supervision, and that is the work being handed to machines first. The mechanism is at cognitive apprenticeship: expertise in cognitive work transfers when the reasoning is made explicit, and the machine's draft removes the occasion for anybody to make it explicit.
- Protect the repetitions that teach. Not all of them. The ones whose feedback is fast and honest, which is the fourth question of the use-or-keep test.
- Ask for the reasoning separately from the document. A junior can now produce a defensible output without having formed a view, and nobody finds out, including the junior. That is synthetic seniority.
- Keep a shared prompt review. Four things on the table: the prompt, the raw output, what the human changed and why, and where it might be wrong.
- Stage the handover deliberately. Medicine is the only profession with explicit machinery for this, in the supervision levels attached to entrustable professional activities. Nobody else has built one, which is a gap rather than a reason to dismiss it.
What to measure#
Four numbers, all cheap to collect and hard to fake, and none of them a score out of anything.
- Override rate. How often human review changes the output. A review that never disagrees is not a review, and this is the number nobody instruments.
- Named ownership. The proportion of consequential decisions where a Hand line names a person.
- The trend in recorded misses. A Head record that stops finding misses usually means somebody stopped looking rather than that the system stopped being wrong.
- Calibration. Of everything called seventy per cent, what share happened. Scored in batches, against forecasts that resolved.
Two things not to measure. A score out of the seventy subskills, which produces a number with no instrument behind it. And anything that turns the five questions into a dashboard metric, since a measure gets optimised and those five do not survive it.
Where the Augmented Mindset sits in all of this#
The Augmented Mindset is what the four stages are developing. It is the one of the seven capabilities that shows up in a record rather than in an impression. A person who has it can produce the receipts, say what they changed in an output and why, name the terrain they were working on, and carry the way they work into an unfamiliar system. A person who has a configuration rather than a mindset can do none of those once the tool changes.
What this programme has not been shown to do#
It has not been run and evaluated as a programme. The four methods were measured separately, in different fields, on outcomes narrower than judgement at work: forecast accuracy, performance on a subsequent exercise, scores on a bias test, skills in a simulator. The step from those to better decisions in an organisation is an inference nobody has tested. The instruments from the book have no outcome evidence at all and are offered as ways of producing a record rather than as interventions. And the ordering is argued rather than trialled: it follows from the fact that review requires a record, which is reasoning, not a finding.
What can be said is that these are the only four methods with measurement behind them, that the alternative on the market has less, and that an organisation running this would at least know what it was doing and be able to show it.
Key sources
- Tannenbaum, S. I. and Cerasoli, C. P. (2013). Do team and individual debriefs enhance performance? A meta-analysis. Human Factors, 55(1).
- Mellers, B. et al. (2014). Psychological strategies for winning a geopolitical forecasting tournament. Psychological Science, 25(5).
- Morewedge, C. K. et al. (2015). Debiasing decisions: improved decision making with a single training intervention.
- Hogarth, R. M., Lejarraga, T. and Soyer, E. (2015). The two settings of kind and wicked learning environments.
- Cook, D. A. et al. (2011). Technology-enhanced simulation for health professions education. JAMA, 306(9).
Related SuperSkills research#
- How is judgement trained
- Hand, Head, Hours, Heart
- The three terrains
- Models of judgement
- Why reskilling programmes mostly fail
Evidence review · SS-2026-342 · Graded against the published rubric
Hirji, R. (2026). How do you build judgement in an organisation?. The SuperSkills evidence base, SS-2026-342. https://thesuperskills.com/research/how-do-you-build-judgement-in-an-organisation. Last reviewed 26 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work