← Research
Research

Human judgement in the age of AI

The agreement is the problem.

Last reviewed: 26 September 2026 · Next review due: 26 September 2027

London Business School, Harvard Business Review, the World Economic Forum, MIT Sloan Management Review, Deloitte and BCG broadly say the same thing. This page names it, says which parts hold, and sets out the disagreement: judgement is being lost by allocation, not by leaders getting worse at thinking.

Questions this page answersAll 996 questions this research covers

Almost everyone writing about human judgement and AI agrees on the same four moves, and the agreement is the problem. This page states the standard account by name, says which parts of it are supported, and sets out where this research parts company with it. The disagreement is one sentence: judgement is being treated as a personal virtue when it is an organisational design property, and the remedies follow from the mistake.

Share this line as a card

The answer, in one line

That AI does prediction or production, judgement is the scarce human complement, the danger is atrophy or automation bias, the remedy is oversight, dissent and critical thinking, and humans still decide what matters.

Share as a card

The standard account, named#

Read enough of it and the structure repeats. AI does prediction or production; judgement is the scarce human complement. The danger is atrophy, automation bias or over-deference. The remedy is oversight, dissent, critical thinking and keeping a human in the loop. It closes on a version of the sentence: humans still decide what matters.

The serious versions are worth reading. Andrew Likierman at London Business School has the most developed account of judgement as a personal capacity, with six elements: knowledge, context, trust, feelings, choice and delivery. Agrawal, Gans and Goldfarb supply the economic frame in which cheap prediction raises the value of judgement. The World Economic Forum has named the category "judgement work". MIT Sloan Management Review has intelligent choice architectures, which is the one line of work in this group asking what happens to decision rights rather than to individuals. Deloitte's 2026 human capital work treats decision-making as a discipline to be built. Harvard Business Review, in August 2026, ran the argument that AI narrows breadth of perception and independence of interpretation, and prescribed structured dissent.

None of that is wrong. It is also, as a body, missing the mechanism.

What the evidence actually supports#

Three things are well enough established to build on, and they are not the three that get quoted.

Combination is not automatically better. Vaccaro, Almaatouq and Malone, across 106 experiments in Nature Human Behaviour in 2024, found human and AI combinations performing on average worse than the better of the two alone, with synergy on creation tasks and not on decision tasks. The default arrangement every organisation is building is the one with the least evidence behind it.

Oversight fails in a specific way. Explanations increase acceptance without improving team accuracy, per Bansal and colleagues in 2021. Automation bias appears in experts and resists training, per Parasuraman and Manzey. The failure is not laziness, it is rational reliance on something that is usually right.

The loss is deferred. The tasks being automated first are the ones through which judgement was built, which is the missing-rungs argument, and the cost appears years after the efficiency does.

Where this research disagrees#

Judgement is not a personal virtue to sharpen. An organisation does not lose judgement because its leaders got worse at thinking. It loses judgement because decisions were reallocated, one reasonable step at a time, until the people nominally accountable could no longer reconstruct how any of it was decided. Sharpening individuals is a training answer to a governance problem, and it has been prescribed for three years without changing anything measurable.

The mechanism is drift, not a skills gap. Nobody decides to hand over judgement. It goes in small allocations that are each defensible: the first draft, the shortlist, the risk rating, the recommendation that arrives already at the top of the list. Drift versus design is the name this research gives it, and it has a dated first publication rather than a claim of novelty.

It is measurable at the organisational level, and almost nobody measures it. Not by asking leaders whether they feel sharp. By asking who can still do the work unaided, how often anyone overrides the system, how long a stop takes and whether a decision could be reconstructed six months later. Those are countable. The instruments are at the noise audit, the Decision Quality Protocol and the delegation boundary map.

The unit of analysis is the decision, not the person. Which decisions may the machines make, who can stop each one, how would you know if one went wrong, and what must your people remain capable of doing so that the humans nominally in charge still are. Four questions, and they are governance questions.

What would change my mind#

A longitudinal study of an organisation that invested heavily in individual judgement training while leaving its allocation of decisions untouched, and improved on a measure that is not self-report. Or a demonstration that appropriate reliance can be produced reliably by design, which would make oversight a control rather than an aspiration. Neither exists today. If either appears it will be dated and recorded at the corrections ledger, alongside the rest.

Where to go next

Every named model in the field, graded with what it does not prove, is at the models of judgement. The definition this research works from is at what is judgement. The evidence on weakening is at does AI weaken human judgement. The position stated as an argument, with the objections, is at judgement.

Essay · SS-2026-310

Cite this page

Hirji, R. (2026). Human judgement in the age of AI. The SuperSkills evidence base, SS-2026-310. https://thesuperskills.com/research/human-judgement-in-the-age-of-ai. Last reviewed 26 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

What is the standard account of human judgement in the age of AI?

That AI does prediction or production, judgement is the scarce human complement, the danger is atrophy or automation bias, the remedy is oversight, dissent and critical thinking, and humans still decide what matters. The serious versions include Andrew Likierman's six elements at London Business School, the prediction-and-judgement frame from Agrawal, Gans and Goldfarb, the World Economic Forum's judgement work, and MIT Sloan Management Review's intelligent choice architectures.

What does the evidence actually support?

Three things. That human and AI combinations are not automatically better, per Vaccaro and colleagues across 106 experiments. That oversight fails in a specific way, with explanations raising acceptance without raising accuracy. And that the loss is deferred, because the tasks being automated first are the ones through which judgement was built.

Why is judgement an organisational problem rather than a personal one?

Because organisations do not lose judgement by their leaders getting worse at thinking. They lose it because decisions were reallocated one reasonable step at a time until nobody accountable could reconstruct how any of it was decided. Sharpening individuals is a training answer to a governance problem, and it has been prescribed for three years without changing anything measurable.

Can organisational judgement be measured?

Parts of it, and cheaply. Who can still do the work unaided, how often anyone overrides the system, how long a stop takes, and whether a decision could be reconstructed six months later. All four are countable, and almost nobody counts them.

What would change this position?

A longitudinal study of an organisation that invested in individual judgement training, left its allocation of decisions untouched and improved on a measure that is not self-report. Or a demonstration that appropriate reliance can be produced reliably by design, which would turn oversight into a control rather than an aspiration. Neither exists today.

In this hub

Definitions

The terms this field uses, defined against their primary sources.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory  ·  Boards  ·  Enquire