← Research
Research · Question

How is judgement trained?

Four things with evidence behind them, several without, and the reason the difference exists.

Last reviewed: 26 September 2026 · Next review due: 26 September 2027

Judgement is the capability everybody now says is scarce, and almost nobody says how it is built. The literature is older and less encouraging than the current discussion suggests. This page sets out what has been measured, what has not, and why the answer turns on feedback rather than on effort.

Question this page answersAll 996 questions this research covers

Four things have real evidence behind them: structured debriefs, calibration training, a single well-designed debiasing exercise, and simulation with graduated responsibility. Most of what is sold as judgement training has none. The reason for the gap is not effort. It is that judgement mostly sits in work whose feedback does not teach.

Share this line as a card

The answer, in one line

Parts of it can. Calibration, the match between confidence and accuracy, improves with about an hour of structured training and the effect has been measured against resolved outcomes.

Share as a card

Start with the discouraging part#

The popular account says expertise comes from accumulated practice. The 2014 meta-analysis of 88 studies found deliberate practice explained about 12 per cent of the variance in performance overall, with the share falling from about 26 per cent in games to about 4 per cent in education and under 1 per cent in the professions. The professional figure deserves the caveat the estate keeps on it: it rests on seven effect sizes, is not statistically significant, and comes from computer programming, military piloting, football refereeing and insurance selling, which is not a sample that speaks for law, medicine or analysis. Ericsson disputed the classification of what counted as deliberate practice at all. What survives the dispute is narrower and still useful: accumulated hours are a weak predictor of professional performance, and the strength of the practice effect falls as the domain becomes less stable and less well specified.

Debiasing is the second discouraging result. Teaching people about cognitive biases, which is the most common form of judgement training on the market, has repeatedly failed to change behaviour. The structured analytic techniques taught in the intelligence community, several of which have been borrowed into business as red teams and devil's advocates, were reviewed by Chang and colleagues in 2018, who found the psychological rationale weak and the accuracy evidence essentially absent. That is absence of testing rather than a demonstrated null. It remains an uncomfortable position for a method in its fourth decade of use.

Why experience alone does not do it#

The mechanism is the feedback, and the clearest statement of it is Hogarth's distinction between kind and wicked learning environments. Where feedback is quick, accurate and drawn from a complete sample, experience builds judgement. Where it is delayed, noisy, or censored by the decisions themselves, experience builds confidence instead, and the person has no internal signal telling them which of the two they have acquired.

Balzer's review of the feedback literature sharpens it further. Outcome feedback, being told whether you were right, produced little improvement in judgement and sometimes made it worse. Cognitive feedback, being shown how you weighted the available information against how the situation actually weights it, produced improvement. Knowing the result is not the same as learning from it, and most professional review supplies only the first.

The four things with evidence#

Structured debriefs. The best-evidenced routine available. Tannenbaum and Cerasoli's meta-analysis of 46 studies found properly conducted debriefs improved subsequent performance by around 20 to 25 per cent, with an average effect size near 0.67. The qualifier matters: the effect depends on structure, on the participants reaching the conclusions themselves rather than being told, and on reviewing how the decision was made rather than how it turned out. See the after action review.

Calibration training. Within the Good Judgment Project's randomised forecasting tournament, about an hour of training in probabilistic reasoning improved accuracy by roughly 6 to 12 per cent against control, replicated across successive years and scored against resolved outcomes. It trains one narrow component, the match between confidence and accuracy, and that component happens to be the one most visible to other people. See calibration training.

One well-designed debiasing exercise. Morewedge and colleagues found in 2015 that a single interactive training game produced large reductions in the biases it targeted, still measurable at eight and twelve weeks, and outperforming an instructional video on the same content. The outcome measures were bias tests rather than decisions with consequences, so the result establishes that debiasing is not uniformly futile rather than that it reaches real work. See do debiasing programmes work.

Simulation with graduated responsibility. The largest literature of the four, almost all of it medical. Cook's meta-analysis of 609 studies found large effects on knowledge, skills and behaviour against no intervention. Alongside it, medicine has built an explicit staging mechanism in entrustable professional activities, where a trainee is permitted progressively less supervision as capability is demonstrated. Whether that transfers outside clinical training is untested. No other profession has come as close to making apprenticeship legible.

Two methods that sit between#

The premortem. Klein's procedure has no trial behind it, and the experimental work underneath it is real: treating an outcome as already certain increased the number and specificity of reasons people produced by about 30 per cent in the original 1989 study. What it improves is reason generation, not decisions. See the premortem.

Making expert reasoning visible. Cognitive apprenticeship names the mechanism by which watching a senior person work transfers anything, which is articulation of the reasoning rather than observation of the output. The trainable version with numbers attached is Klein's ShadowBox method, where a trainee makes the same calls an expert made and then compares their cue selection against the expert panel's. The reported gains, about 28 per cent with 59 Marines and 21 per cent with 30 Army officers, come from small evaluations run by the method's developers. See cognitive apprenticeship.

What changes when AI does the first draft#

Everything above depends on a loop: you commit to an answer, something tells you how it went, and you adjust. AI intervenes at the first step. If the machine produces the answer and the person edits it, the person never commits, so there is nothing for the feedback to attach to. Lee and colleagues at Microsoft Research, surveying 319 knowledge workers about 936 real uses, found that higher confidence in the tool predicted less critical thinking, and that effort shifted from producing work to checking it.

The practical response that has been published is narrow and comes from clinical medicine: keep a quota of unaided work, benchmark how often the human agrees with the machine as a way of detecting drift into deference, delay access to the tool until baseline competence exists, and run scenarios where the tool is wrong. None of that has been trialled. It is the only published attempt to specify how to keep the loop alive, and organisations wanting to do something have little else to work from.

What this page does not establish#

That any of these methods builds judgement in general. Each was measured on something narrower: forecast accuracy, performance on a subsequent exercise, scores on a bias test, skills in a simulator. The step from those to judgement at work is an inference that nobody has tested. The four methods also come from different fields and have never been compared against each other. What can be said is that they are the only four with measurement behind them, and that the many programmes without measurement are not entitled to assume they work.

Key sources

Evidence review · SS-2026-318 · Graded against the published rubric

Cite this page

Hirji, R. (2026). How is judgement trained?. The SuperSkills evidence base, SS-2026-318. https://thesuperskills.com/research/how-is-judgement-trained. Last reviewed 26 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

Can judgement be trained?

Parts of it can. Calibration, the match between confidence and accuracy, improves with about an hour of structured training and the effect has been measured against resolved outcomes. Structured debriefs improve subsequent performance by around a fifth across 46 studies. Simulation with graduated responsibility is the best-evidenced delivery structure. Broad claims about training judgement in general have far less behind them.

Does deliberate practice build judgement?

Less than the popular account suggests. The 2014 meta-analysis found deliberate practice explained about 12 per cent of variance in performance overall and under 1 per cent in the professions, though that professional figure rests on seven effect sizes, is not statistically significant, and comes from four occupations that do not include law, medicine or analysis. The definition of deliberate practice used in those studies is itself disputed.

Does debiasing training work?

Mostly not, with one well-measured exception. Telling people about biases changes little. A single interactive training game in Morewedge's 2015 experiments produced large reductions in the targeted biases that were still present at twelve weeks. The outcome measures were bias tests rather than real decisions.

Why is judgement hard to learn from experience?

Because most professional judgement sits in what Hogarth calls a wicked learning environment: the feedback is delayed, partial, or censored by the decisions themselves. Experience in such a setting reliably produces confidence and does not reliably produce accuracy.

In this hub

Definitions

The terms this field uses, defined against their primary sources.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory  ·  Boards  ·  Enquire