Groups can be better or worse than their members, and which one happens depends on structure rather than on the people. The one reliable improvement requires nobody to think better: collect judgements independently, then combine them.
The answer, in one line
Only under conditions. Aggregating independent judgements reliably reduces error, and it works without anybody becoming less biased.
Definition#
Collective judgement: judgement produced by a group. It improves on the individual reliably where judgements are made independently before being combined; where they are formed in discussion, the group inherits the first framing offered and can perform worse than its average member.
What works#
Aggregating independent judgements reduces error, and it does so without anybody becoming less biased. Larrick's review of the debiasing literature identifies it as one of the few interventions that transfers, alongside incentives aligned to accuracy and the instruction to consider the opposite. It is structural rather than cognitive, so it survives contact with ordinary people under time pressure.
The condition is independence, and independence is destroyed casually. A judgement offered after somebody senior has spoken is not independent. A judgement formed during a discussion of the options is not independent. A judgement made by five people who all read the same briefing note is only partly independent, so the practical instruction is to collect the numbers privately, before the meeting, and to discuss afterwards.
What has almost no evidence#
Almost everything organisations actually do. Red teams, devil's advocates, competing hypotheses and the rest of the structured analytic techniques developed in intelligence work were reviewed by Chang, Berdini, Mandel and Tetlock in 2018, who found the psychological rationale weak and evidence of improved accuracy essentially absent. That is absence of testing rather than a demonstrated null, and it remains an uncomfortable position for methods in their fourth decade of institutional use.
The premortem is the strongest of the family and its support is for generating more and better specified reasons rather than for better outcomes. What all of these share is a sound social mechanism, lowering the cost of stating a doubt, and no measurement of whether the doubts that surface are the ones that mattered.
Forecasting as the exception#
Where groups have been measured properly is forecasting, because the questions resolve. The Good Judgment work found that accuracy improved with brief probabilistic training, that teams outperformed individuals, and that the best forecasters clustered. The conditions that made that measurable, dated questions with unambiguous answers, are the conditions most organisational judgement lacks, so the result is real and hard to transfer. See calibration training.
What AI does to a group#
It supplies a shared starting point, and a shared starting point is the condition under which aggregation stops working. If six people begin from the same generated draft, their judgements are correlated before anybody speaks, and the agreement that follows measures the model rather than the group. The apparent consensus is the most dangerous form, because it looks like the independent convergence that aggregation is built on.
The remedy is the same as the general one and needs stating explicitly: form a view before opening the tool, record it, and treat the model's output as one more independent input rather than as the frame. The related pattern is at does AI make everyone think alike.
The practical version#
- Collect judgements privately, with a number, before the meeting.
- Show the spread before discussing the answer. The spread is the finding; a wide one on a decision everybody assumed was settled is the thing a noise audit would have found.
- Discuss afterwards, and re-collect.
- Have the most junior person speak first, which is free and reverses the usual contamination.
What this does not establish#
That groups run this way make better decisions in organisations. The aggregation evidence is strongest in settings with numerical estimates and resolvable answers, and its extension to a leadership team deciding a strategy is an inference. Nothing here shows that the dissent techniques fail, only that they have not been shown to work.
Key sources
- Chang, W., Berdini, E., Mandel, D. R. and Tetlock, P. E. (2018). Restructuring structured analytic techniques in intelligence.
- Mellers, B. et al. (2014). Psychological strategies for winning a geopolitical forecasting tournament.
Related SuperSkills research#
Explainer · SS-2026-345 · Graded against the published rubric
Hirji, R. (2026). What is collective judgement?. The SuperSkills evidence base, SS-2026-345. https://thesuperskills.com/research/what-is-collective-judgement. Last reviewed 26 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work