Calibration is the match between how confident you are and how often you are right. It is the one component of judgement that has been trained and measured against real outcomes. About an hour of structured training improved forecasting accuracy by roughly 6 to 12 per cent in a randomised tournament, replicated across successive years.
The answer, in one line
The match between how confident a person is and how often they are right. A well calibrated person is right about 70 per cent of the time when they say they are 70 per cent sure.
Definition#
Calibration training: structured training in probabilistic reasoning, usually covering comparison classes, base rates, incremental updating and review of resolved forecasts, aimed at bringing stated confidence into line with actual accuracy.
Calibration is not accuracy#
A well calibrated person is right about seven times in ten when they say they are seventy per cent sure. They may know very little, in which case their confidence will be low and correctly low. Accuracy is about being right; calibration is about knowing how likely you are to be right. The two come apart constantly, and the second is the one other people rely on, because a colleague who says they are certain is making a claim you will act on.
Overconfidence is the usual direction of the error, and it grows with expertise in fields where feedback is poor. That is the same mechanism described at kind and wicked learning environments: confidence accumulates with experience whether or not accuracy does.
What the tournament showed#
The Good Judgment Project ran within a multi-year forecasting tournament funded by the US intelligence community, with participants randomly assigned to conditions and forecasts scored against events that resolved. A short training module in probabilistic reasoning, taking about an hour, improved accuracy by roughly 6 to 12 per cent against control, and the effect held across successive years of the tournament.
The content is unglamorous: start from a comparison class and its base rate rather than from the specifics of the case, update in increments rather than in leaps, distinguish the question you were asked from the question you find interesting, and review your resolved forecasts. Earlier work by Lichtenstein and Fischhoff in 1980 had already shown that calibration improves with feedback about calibration specifically, with most of the gain arriving early.
Why this result is smaller than it sounds#
Because the tournament scored questions with dates and resolvable answers, and almost nothing in professional life has those. Whether a hire works out, whether a strategy was right, whether a restructure helped: none of these resolve cleanly, on a date, against a stated forecast. The result establishes that a narrow component of judgement is trainable in a setting built for measurement. It does not establish that judgement at work improves.
What does appear to transfer is the habit rather than the skill. Stating a probability instead of an adjective, writing it down, and looking back at it is available in any job, and it converts a corner of a wicked environment into a kind one.
Doing it without a tournament#
- Attach a number to predictions that will resolve within months, and record it. A decision journal is the container.
- Score in batches. Of everything you called 70 per cent, what share happened?
- Start from the base rate for cases like this one before considering what makes this one different.
- Keep the questions boring and resolvable. The point is the scoring, not the subject.
Calibration and machines#
Two calibrations are now in play. A model reports confidence that may not track its accuracy, and a person forms confidence in the model that may not track its reliability either. The second is the subject of appropriate reliance, and the research there is not encouraging: neither explanations nor confidence scores have reliably produced calibrated trust. Training a person's own calibration does not fix that. It does make the person's contribution legible, which is the part a colleague or a board can actually use.
What this does not establish#
That better calibrated professionals make better decisions. The tournament measured forecast accuracy, not decision quality, and the two are related without being the same. Nor is there evidence that the effect persists without continued scoring; the earlier calibration work found improvement was largely confined to the trained task.
Key sources
- Mellers, B. et al. (2014). Psychological strategies for winning a geopolitical forecasting tournament. Psychological Science, 25(5).
- Lichtenstein, S. and Fischhoff, B. (1980). Training for calibration. Organizational Behavior and Human Performance, 26(2).
Related SuperSkills research#
Explainer · SS-2026-324 · Graded against the published rubric
Hirji, R. (2026). What is calibration training?. The SuperSkills evidence base, SS-2026-324. https://thesuperskills.com/research/what-is-calibration-training. Last reviewed 26 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work