- Is human judgement a skill or an organisational property?
- What do business schools say about AI and judgement?
Almost everyone writing about human judgement and AI agrees on the same four moves, and the agreement is the problem. This page states the standard account by name, says which parts of it are supported, and sets out where this research parts company with it. The disagreement is one sentence: judgement is being treated as a personal virtue when it is an organisational design property, and the remedies follow from the mistake.
The answer, in one line
That AI does prediction or production, judgement is the scarce human complement, the danger is atrophy or automation bias, the remedy is oversight, dissent and critical thinking, and humans still decide what matters.
The standard account, named#
Read enough of it and the structure repeats. AI does prediction or production; judgement is the scarce human complement. The danger is atrophy, automation bias or over-deference. The remedy is oversight, dissent, critical thinking and keeping a human in the loop. It closes on a version of the sentence: humans still decide what matters.
The serious versions are worth reading. Andrew Likierman at London Business School has the most developed account of judgement as a personal capacity, with six elements: knowledge, context, trust, feelings, choice and delivery. Agrawal, Gans and Goldfarb supply the economic frame in which cheap prediction raises the value of judgement. The World Economic Forum has named the category "judgement work". MIT Sloan Management Review has intelligent choice architectures, which is the one line of work in this group asking what happens to decision rights rather than to individuals. Deloitte's 2026 human capital work treats decision-making as a discipline to be built. Harvard Business Review, in August 2026, ran the argument that AI narrows breadth of perception and independence of interpretation, and prescribed structured dissent.
None of that is wrong. It is also, as a body, missing the mechanism.
What the evidence actually supports#
Three things are well enough established to build on, and they are not the three that get quoted.
Combination is not automatically better. Vaccaro, Almaatouq and Malone, across 106 experiments in Nature Human Behaviour in 2024, found human and AI combinations performing on average worse than the better of the two alone, with synergy on creation tasks and not on decision tasks. The default arrangement every organisation is building is the one with the least evidence behind it.
Oversight fails in a specific way. Explanations increase acceptance without improving team accuracy, per Bansal and colleagues in 2021. Automation bias appears in experts and resists training, per Parasuraman and Manzey. The failure is not laziness, it is rational reliance on something that is usually right.
The loss is deferred. The tasks being automated first are the ones through which judgement was built, which is the missing-rungs argument, and the cost appears years after the efficiency does.
Where this research disagrees#
Judgement is not a personal virtue to sharpen. An organisation does not lose judgement because its leaders got worse at thinking. It loses judgement because decisions were reallocated, one reasonable step at a time, until the people nominally accountable could no longer reconstruct how any of it was decided. Sharpening individuals is a training answer to a governance problem, and it has been prescribed for three years without changing anything measurable.
The mechanism is drift, not a skills gap. Nobody decides to hand over judgement. It goes in small allocations that are each defensible: the first draft, the shortlist, the risk rating, the recommendation that arrives already at the top of the list. Drift versus design is the name this research gives it, and it has a dated first publication rather than a claim of novelty.
It is measurable at the organisational level, and almost nobody measures it. Not by asking leaders whether they feel sharp. By asking who can still do the work unaided, how often anyone overrides the system, how long a stop takes and whether a decision could be reconstructed six months later. Those are countable. The instruments are at the noise audit, the Decision Quality Protocol and the delegation boundary map.
The unit of analysis is the decision, not the person. Which decisions may the machines make, who can stop each one, how would you know if one went wrong, and what must your people remain capable of doing so that the humans nominally in charge still are. Four questions, and they are governance questions.
What would change my mind#
A longitudinal study of an organisation that invested heavily in individual judgement training while leaving its allocation of decisions untouched, and improved on a measure that is not self-report. Or a demonstration that appropriate reliance can be produced reliably by design, which would make oversight a control rather than an aspiration. Neither exists today. If either appears it will be dated and recorded at the corrections ledger, alongside the rest.
Where to go next
Every named model in the field, graded with what it does not prove, is at the models of judgement. The definition this research works from is at what is judgement. The evidence on weakening is at does AI weaken human judgement. The position stated as an argument, with the objections, is at judgement.
Essay · SS-2026-310
Hirji, R. (2026). Human judgement in the age of AI. The SuperSkills evidence base, SS-2026-310. https://thesuperskills.com/research/human-judgement-in-the-age-of-ai. Last reviewed 26 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work