← Research
Research

AI and human judgement

AI can raise the quality of an output while weakening the capability that would have produced it. That gap is the story of the age.

Last reviewed: 25 August 2026

Does AI reduce human judgement? Not on its own, and not for everyone. But the risk is real and specific, and it is the central question in Rahim Hirji's work and in SuperSkills (Kogan Page, 2026). This page separates what the evidence shows from how to read it.

Used well, AI sharpens judgement. Used by default, it erodes it. The evidence does not show that artificial intelligence makes people less intelligent. It shows something more precise, and more troubling: AI can raise the quality of a single output while weakening the underlying capability that would have produced that output unaided. A knowledge worker who leans on a model to reason, draft and decide gets a better result today and does the thinking less often. Over months, the repetitions that build judgement are quietly transferred to the machine. The output improves. The person does not. Whether AI reduces your judgement is therefore not a question about the technology. It is a question about how deliberately you design the relationship between the human and the machine. Most organisations do not design it at all. They drift.

What the evidence shows

Start with the mechanism, because it predates AI. Psychologists call it cognitive offloading: using an external tool to reduce the mental demand of a task. Risko and Gilbert, in a 2016 review, showed that we offload not only when a task is genuinely hard but when we judge it to be hard, and that this judgement is often wrong. We hand work to tools we did not need to, and we lose the practice we would otherwise have had.

Two long-running lines of research show where that leads. Sparrow, Liu and Wegner, writing in Science in 2011, found the Google effect: when people expect information to remain available, they remember where to find it rather than the thing itself. Dahmani and Bohbot, in 2020, found that habitual users of satellite navigation had worse spatial memory when asked to navigate unaided, and that heavier GPS use over the following three years was associated with a steeper decline still. The pattern is consistent: when a capability is reliably performed by an external system, the human capability for it tends to weaken. Memory and navigation were the first to go. Reasoning and judgement are simply the next functions to be offloaded, and they are more central to professional work than either.

The early workplace evidence on generative AI fits the pattern. A 2025 study by Microsoft Research and Carnegie Mellon surveyed 319 knowledge workers about 936 real uses of AI in their jobs. It found that higher confidence in the tool was associated with less critical thinking, and that the nature of the thinking shifts: from gathering information to verifying the machine's output, from solving the problem to integrating the answer, from doing the task to supervising it. Critical thinking does not disappear. It moves, and it thins. A separate 2025 study by Michael Gerlich, across 666 participants, found a negative correlation between frequent AI use and critical-thinking scores, with cognitive offloading as the mechanism in between, and the effect strongest among the youngest users, who have offloaded earliest.

The most striking single result comes from the field. In 2023, Boston Consulting Group and researchers at Harvard, MIT and Wharton, among them Karim Lakhani, ran a controlled experiment with 758 consultants using GPT-4, and named what they found the jagged technological frontier. Inside the frontier, on tasks the model handled well, AI-assisted consultants were dramatically better and faster. Outside it, on a task designed to sit just beyond the model's competence, consultants using AI did worse than those with no AI at all, because they accepted confident output they should have questioned. That is automation bias, and it is not new. Parasuraman and Manzey, reviewing decades of work in aviation, medicine and the military in 2010, showed that the tendency to under-question automated advice appears in novices and experts alike, cannot be trained away, and worsens under load. What was once a cockpit problem is now an office problem.

None of this means AI makes work worse. It plainly makes a great deal of it better. Brynjolfsson, Li and Raymond, studying 5,179 customer-support agents, found that access to an AI assistant raised productivity by fourteen percent on average, and by thirty-four percent for the newest and least experienced staff, while barely moving the most skilled. AI transfers the patterns of expert workers to inexperienced ones. That is a real and valuable gain. It also raises the question this page exists to ask: if the tool carries the novice to expert-looking output, what happens to the experience through which the novice was supposed to become an expert?

The honest summary of the evidence is not that AI makes us stupid. It is narrower and harder to dismiss: we get worse at what we stop practising, and AI is very good at letting us stop practising the exact thing, judgement, that professional value rests on.

Where the evidence remains uncertain

The direction of the risk is well supported. Its size is not, and honesty about that matters more than a confident headline. Much of the strongest recent work is self-reported or correlational. The Microsoft and Carnegie Mellon study asked workers to describe their own thinking; people who already think differently may use AI differently, and a survey cannot separate the two. Gerlich's study shows a correlation, not proof that AI use causes the decline rather than accompanying it.

A 2025 study from the MIT Media Lab, using EEG to compare people writing essays with a language model, a search engine, or unaided, found the weakest brain connectivity and the lowest sense of ownership in the AI group, and called the effect cognitive debt. It is a striking result, but it rests on 54 participants, it is a preprint, and its own later commentators flag sample size and reproducibility. It is suggestive, not settled. Treat anyone who cites it as conclusive with the caution the study's own authors would want.

The productivity studies, by contrast, are not in doubt that output and speed rise, at least in the short term. The capability question operates on a longer clock, over the years across which judgement is actually built or lost, and that is precisely the horizon the workplace studies have not yet had time to measure. The navigation and memory research is the closest long-run analogue we have, and it points one way, but reasoning is not spatial memory, and the analogy should carry weight without carrying certainty. The responsible position is this: the risk is real enough, and slow enough to be invisible, that the time to design against it is now, not after a decade of proof arrives.

The SuperSkills interpretation

Hold the two facts together, because almost every argument about AI picks one and drops the other. AI can improve the quality of an individual output and weaken the capability that would have produced that output unaided. The optimists cite the first and call the worry technophobia. The pessimists cite the second and call the gains a mirage. Both are describing the same event from opposite ends. The output goes up. The practice goes down. Whether that is progress or erosion depends entirely on what you do about the practice.

This reframes the question that dominates the public conversation. The popular fear is that AI will take your job. The more immediate risk is that AI will take your judgement, and the two move at very different speeds. Jobs change slowly and visibly, through restructures and headcount, where someone at least has to decide. Judgement erodes quietly and by default, one delegated decision at a time, with nobody choosing it and nobody able to point to the moment it happened. That is why I describe it as a drift problem rather than a decision. No leadership team sets out to hollow out its own people. It happens because the tool arrives faster than the design. The whole of drift versus design is the choice to decide in advance where human judgement has to remain, rather than letting the default settle it for you.

Underneath the drift are four mechanisms, each of which has its own page in this research. The missed reps are the repetitions through which judgement is built, now handed to the machine, so the work ships but the practice never happens. The missing rungs are the junior tasks that used to carry people up to senior judgement, removed by automation before anyone noticed they were load-bearing. Synthetic seniority is the result at the level of the individual: output that looks like judgement without the judgement underneath. And the accumulated organisational version is what I call capability debt: the loss of human knowledge, skill and judgement that builds up when an organisation automates work faster than it redesigns how people learn through doing. Capability debt is invisible on any dashboard, because the outputs still look fine, right up until the moment a decision arrives that the AI cannot make and no human in the room has been trained to.

So the answer to whether AI reduces human judgement is this. AI does not reduce judgement on its own. It reduces the demand for judgement, and demand is what builds and maintains it. Remove the demand without redesigning how people learn, and judgement follows memory and navigation down the same slope. Keep the demand deliberately, and AI becomes what it should be: a tool that raises the floor of what people can produce while you protect the ceiling of what they can think.

What leaders should do

Decide where human judgement must remain, and decide it in advance. This is the practical core of drift versus design. The question is not whether to adopt AI, which is settled, but which decisions a human has to own, and why, and to make that list explicit rather than letting each case be resolved by whoever is busiest that afternoon. An organisation that has never written the list has already answered by default.

Protect the repetitions. If juniors never do the task the AI now does, they never build the judgement the senior role will demand of them. That does not mean banning the tool, which is neither enforceable nor wise. It means keeping deliberate practice in the system on purpose: some work done unaided so the capability is exercised, some AI-assisted work annotated so the person can say what they prompted, what the model gave them and what they changed, and some judgement tested directly rather than inferred from a polished output.

Treat verification as real work, not residue. The jagged-frontier result is the warning every leader should keep in view: the people who trusted AI outside its competence did worse than people with no AI at all. Verification is not a rubber stamp at the end of a process; it is the judgement layer, and it is often harder than production. Staff it, train it and value it accordingly, or you will pay least for the work you depend on most.

Measure capability, not only output. Output quality has quietly stopped being a reliable proxy for the capability of the person who submitted it, which breaks the assumption most promotion and performance systems are built on. Borrow from the professions that solved this long ago: medicine and aviation test judgement directly, through live decisions, simulation and oral examination, rather than trusting that good work implies a capable person. And watch the development curve, not just the productivity curve. The gains Brynjolfsson found flow to novices because AI hands them expert patterns; make sure it is teaching them, not merely carrying them.

The market is already moving this way. The World Economic Forum's 2025 Future of Jobs report names analytical thinking as the single most valued core skill among employers, and skills gaps as the biggest barrier to transformation over the next five years. The organisations that come out ahead will not be the ones that adopted AI fastest. They will be the ones that decided, deliberately, where human judgement belongs, and built the practice to keep it.

Key research and primary sources

Where a claim matters, go to the study rather than the article reporting it. These are the primary sources behind this page.

Related SuperSkills research

The concepts in this page are developed in their own right elsewhere: drift versus design, the missed reps, the missing rungs, synthetic seniority, the verifier's discount and decision quality. Together they make up the SuperSkills account of what AI adoption does to human capability.

About this research

Rahim Hirji is the author of SuperSkills: The Seven Human Skills for the Age of AI (Kogan Page, 2026) and the founder of The SuperSkills Intelligence Company. This work draws on research across more than 200 organisations in 30 countries over seven years. The page is written to a deliberate rule: findings are attributed to the studies that produced them, and kept separate from the interpretation, which is the author's. The named concepts, drift versus design, the missed reps, the missing rungs, synthetic seniority and capability debt, are part of the SuperSkills lexicon. This is a living reference, reviewed and updated as significant new evidence appears, rather than a dated article left to stand.

Cite this

Hirji, R. (2026). AI and human judgement. The SuperSkills Intelligence Company. Last reviewed 25 August 2026. thesuperskills.com/research/ai-and-human-judgement

The work

Where the writing comes from.

These essays draw on research across more than 200 organisations in 30 countries. See the wider body of work, or bring it into your organisation.

All research →