← Research
Research · Definition

What is collective judgement?

The one reliable improvement is structural, and it requires nobody to think better.

Last reviewed: 26 September 2026 · Next review due: 26 September 2027

Most of what organisations do about group judgement is a technique for producing dissent. The technique literature is thin. The aggregation literature is not, and it asks for something organisations find harder than a workshop.

Questions this page answersAll 996 questions this research covers

Groups can be better or worse than their members, and which one happens depends on structure rather than on the people. The one reliable improvement requires nobody to think better: collect judgements independently, then combine them.

The answer, in one line

Only under conditions. Aggregating independent judgements reliably reduces error, and it works without anybody becoming less biased.

Share as a card

Definition#

Collective judgement: judgement produced by a group. It improves on the individual reliably where judgements are made independently before being combined; where they are formed in discussion, the group inherits the first framing offered and can perform worse than its average member.

Share this definition as a card

What works#

Aggregating independent judgements reduces error, and it does so without anybody becoming less biased. Larrick's review of the debiasing literature identifies it as one of the few interventions that transfers, alongside incentives aligned to accuracy and the instruction to consider the opposite. It is structural rather than cognitive, so it survives contact with ordinary people under time pressure.

The condition is independence, and independence is destroyed casually. A judgement offered after somebody senior has spoken is not independent. A judgement formed during a discussion of the options is not independent. A judgement made by five people who all read the same briefing note is only partly independent, so the practical instruction is to collect the numbers privately, before the meeting, and to discuss afterwards.

What has almost no evidence#

Almost everything organisations actually do. Red teams, devil's advocates, competing hypotheses and the rest of the structured analytic techniques developed in intelligence work were reviewed by Chang, Berdini, Mandel and Tetlock in 2018, who found the psychological rationale weak and evidence of improved accuracy essentially absent. That is absence of testing rather than a demonstrated null, and it remains an uncomfortable position for methods in their fourth decade of institutional use.

The premortem is the strongest of the family and its support is for generating more and better specified reasons rather than for better outcomes. What all of these share is a sound social mechanism, lowering the cost of stating a doubt, and no measurement of whether the doubts that surface are the ones that mattered.

Forecasting as the exception#

Where groups have been measured properly is forecasting, because the questions resolve. The Good Judgment work found that accuracy improved with brief probabilistic training, that teams outperformed individuals, and that the best forecasters clustered. The conditions that made that measurable, dated questions with unambiguous answers, are the conditions most organisational judgement lacks, so the result is real and hard to transfer. See calibration training.

What AI does to a group#

It supplies a shared starting point, and a shared starting point is the condition under which aggregation stops working. If six people begin from the same generated draft, their judgements are correlated before anybody speaks, and the agreement that follows measures the model rather than the group. The apparent consensus is the most dangerous form, because it looks like the independent convergence that aggregation is built on.

The remedy is the same as the general one and needs stating explicitly: form a view before opening the tool, record it, and treat the model's output as one more independent input rather than as the frame. The related pattern is at does AI make everyone think alike.

The practical version#

What this does not establish#

That groups run this way make better decisions in organisations. The aggregation evidence is strongest in settings with numerical estimates and resolvable answers, and its extension to a leadership team deciding a strategy is an inference. Nothing here shows that the dissent techniques fail, only that they have not been shown to work.

Key sources

Explainer · SS-2026-345 · Graded against the published rubric

Cite this page

Hirji, R. (2026). What is collective judgement?. The SuperSkills evidence base, SS-2026-345. https://thesuperskills.com/research/what-is-collective-judgement. Last reviewed 26 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

Do groups make better judgements than individuals?

Only under conditions. Aggregating independent judgements reliably reduces error, and it works without anybody becoming less biased. Judgements formed in discussion do not aggregate: the group inherits the framing of whoever spoke first and the confidence of whoever spoke loudest, and can perform worse than its average member.

Do red teams and devil's advocates work?

There is very little evidence either way. Chang, Berdini, Mandel and Tetlock reviewed the structured analytic techniques taught in intelligence work in 2018 and found the psychological rationale weak and accuracy evidence essentially absent. That is a gap in testing rather than a demonstrated failure, and an uncomfortable position for methods in their fourth decade of use.

What is the practical version?

Collect judgements privately before the meeting, with a number attached. Discuss afterwards. The order is the intervention: cheap, unglamorous and routinely reversed.

What does AI do to group judgement?

It supplies a shared starting point, which is the condition under which aggregation stops working. If everybody in the room began from the same generated draft, their judgements are no longer independent, and the apparent agreement that follows is a measurement of the model rather than of the group.

In this hub

Definitions

The terms this field uses, defined against their primary sources.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory  ·  Boards  ·  Enquire