A noise audit measures how much two people in the same role disagree when they judge the same case. It is the only instrument in the judgement literature that gives an organisation a number before anybody argues about opinions. Kahneman, Sibony and Sunstein set it out in Noise: A Flaw in Human Judgment (2021). It is worth knowing precisely, because it is the closest thing this field has to a measurement, and because what it does not fix matters as much as what it does.
The answer, in one line
A noise audit is a controlled exercise in which several professionals independently judge the same set of real cases, so the variation between them can be measured.
Definition#
A noise audit: a controlled exercise in which several professionals independently judge the same set of realistic cases, so that the variation between them can be measured. The variation is the noise. Bias is the shared error that moves every judgement in one direction; noise is the scatter, and an organisation usually has no idea how large its own is until it looks.
Noise is not bias, and the difference decides what you do next#
Bias and noise are both error and they call for opposite remedies. Bias is a systematic tilt: everybody in the department is too optimistic about delivery dates. Noise is inconsistency: two underwriters, two clinicians, two hiring managers, two grant assessors reach different conclusions on identical facts, and which one you got was a lottery.
The book separates the scatter into parts that behave differently. Level noise is the difference in where individuals set their general severity. Pattern noise is the stable idiosyncrasy of a person's reactions to particular cases. Occasion noise is the same person differing from themselves, on a different day, in a different mood, after lunch. A single number hides all three, so an audit reports the spread rather than an average.
Why it matters more once AI is in the workflow#
A model is quiet by construction. Give it the same input twice and it returns something close to the same output, so replacing a noisy human process with a model can look like an improvement on every dashboard. Consistency is not accuracy. A system that is quiet and wrong is worse than one that is noisy and wrong, because the error is now uniform, defensible and invisible, and nobody in the chain will notice it from the variation.
The practical consequence for anyone putting a model into a judgement process: measure the noise before you automate, because that measurement is the only honest baseline you will get. Afterwards you are comparing the new system against a memory of the old one.
How it is run#
- Take real cases, not vignettes. Between ten and twenty files the organisation has actually handled, chosen to span the ordinary range rather than the interesting extremes.
- Judge them independently. No discussion, no shared meeting, no visibility of anybody else's answer. Discussion produces agreement without producing accuracy.
- Record the judgement in the unit the job uses. A number, a band, a recommendation, a yes or no. If the work has no expressible unit, that is itself a finding.
- Report the spread rather than the mean. The mean is what managers ask for and it is the one statistic that cannot show noise.
- Ask people to predict the spread first. The gap between the predicted disagreement and the measured disagreement is usually the part that changes behaviour.
What it does not establish#
Reducing noise does not raise accuracy on its own. A process made perfectly consistent around a biased view is now reliably wrong, and it will produce cleaner audit trails while doing it. The audit measures disagreement, not correctness, and correctness needs outcomes the organisation usually does not collect.
The remedies the book proposes, decision hygiene, structured assessments, aggregation of independent views, are argued from the same logic rather than demonstrated by trials in the organisations that adopt them. The strongest empirical support underneath the argument is old and sits in the clinical prediction literature: see clinical versus actuarial judgement, where mechanical consistency beats expert judgement by a modest and frequently overstated margin.
One further limit, for this estate's purposes: a noise audit measures judgement where the work produces a comparable unit. Strategy, hiring at the top, research direction and most board decisions do not, and the literature has no instrument for those at all.
Where it sits in this research#
It is the measurement half of the argument on this site. The Decision Quality Protocol asks who owns a decision and what quality control looks like; a noise audit is the cheapest way to answer the second of those with evidence rather than assertion. On what the models of judgement say together, and where they contradict each other, see the models of judgement.
Key sources
- Kahneman, D., Sibony, O. and Sunstein, C. R. (2021). Noise: A Flaw in Human Judgment. Little, Brown Spark.
- Kahneman, D., Lovallo, D. and Sibony, O. (2019). A Structured Approach to Strategic Decisions. MIT Sloan Management Review. The Mediating Assessments Protocol.
- Grove, W. M., Zald, D. H., Lebow, B. S., Snitz, B. E. and Nelson, C. (2000). Clinical versus mechanical prediction: a meta-analysis. Psychological Assessment, 12(1).
Explainer · SS-2026-314 · Graded against the published rubric
Hirji, R. (2026). What is a noise audit?. The SuperSkills evidence base, SS-2026-314. https://thesuperskills.com/research/what-is-a-noise-audit. Last reviewed 26 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work