Designed well, AI improves the decisions a system produces. Adopted by default, it removes the occasions on which human judgement was built. That single distinction explains why the evidence on this question looks contradictory: AI can raise the quality of a particular output while weakening the capability that would have produced that output unaided. Both things are true at once, they are measured at different moments, and almost every public argument about AI picks one and drops the other.
The short answer#
Not on its own, and not for everyone. AI does not reach into a person and remove their reasoning. What it removes is the demand for reasoning. Demand is what builds judgement in the first place and what maintains it afterwards. Remove the demand without redesigning how people learn, and there is no reason to expect judgement to behave better than the functions we have already handed over. Keep the demand deliberately, and AI raises the floor of what people can produce while the ceiling of what they can think is protected.
Three findings anchor that answer, and not one of them comes from this research.
- A meta-analysis of 106 experimental studies found that human and AI combinations performed significantly worse on average than the better of the human alone or the AI alone, with the losses concentrated in decision-making tasks (Vaccaro, Almaatouq and Malone, 2024).
- Endoscopists averaging twenty-eight years of experience saw their unassisted adenoma detection rate fall from 28.4 to 22.4 per cent after routine exposure to an AI detection tool (Budzyń et al., 2025).
- Students given unrestricted access to GPT-4 scored 48 per cent higher while they had it, and 17 per cent lower than a control group who never had it once it was taken away (Bastani et al., 2025).
The first says the pairing is not automatically good. The second says exposure can leave an experienced professional worse than before. The third says the gain and the loss can be the same event, observed at two different moments. Everything below is an attempt to say precisely when each applies.
What gets removed is the demand#
Start with the mechanism, because it long predates AI. Psychologists call it cognitive offloading: using an external tool to reduce the mental demand of a task. Risko and Gilbert, reviewing the experimental literature in 2016, showed that people offload not only when a task is hard but when they judge it to be hard, and that this metacognitive judgement is frequently mistaken. We hand away work we did not need to hand away, and we lose the practice we would otherwise have had.
Two findings show where that leads. Sparrow, Liu and Wegner, in four laboratory experiments published in Science in 2011, found what became known as the Google effect: when people expect information to remain available, they remember where to find it rather than the thing itself. Dahmani and Bohbot, in 2020, found that habitual satellite navigation users had worse spatial memory when navigating unaided, and that heavier use across the following three years was associated with a steeper decline still.
Neither of those is about judgement, and the distinction deserves more weight than it usually gets. Spatial memory is not reasoning. Sparrow shows a change in what gets encoded, and explicitly does not establish that total memory capability declines or that the trade is a net loss. Dahmani and Bohbot is the only one of the two with a long-run design. That study is correlational. What the pair establish is narrower than the use they are commonly put to: where an external system reliably performs a function, what people retain of that function changes, and the change is gradual enough that nobody notices the day it starts. Memory and navigation were the first functions cheap enough to offload. Reasoning and judgement are the next, and they sit closer to the centre of professional work than either.
This is the claim the rest of the page tests. Here it is in the plainest form I can put it, so that it can be disagreed with: AI reduces the number of occasions on which a person has to exercise judgement, and those occasions were the mechanism by which judgement was acquired and kept.
What happens when humans and AI decide together#
The most useful result for this question is also among the least quoted. Vaccaro, Almaatouq and Malone published a preregistered systematic review and meta-analysis in Nature Human Behaviour in 2024, covering 106 experimental studies and 370 effect sizes. Human and AI combinations performed significantly worse on average than the better of human alone or AI alone, at a Hedges' g of minus 0.23. The losses concentrated in decision-making. The gains, where they appeared, concentrated in content creation.
Read the shape of it rather than the headline. Pairing gained where humans were already better than the AI, and lost where the AI was already better than humans. The combination anchors on the human rather than selecting whichever party is stronger. Two limits travel with the result. The comparison is against an oracle who always picks the better performer, which nobody can do in advance. And the studies were published between January 2020 and June 2023, so the meta-analysis predates the current generation of frontier models. It is an argument against assuming the pairing is free, rather than proof that it cannot be made to pay.
The best-known field experiment says the same thing in a different register. Dell'Acqua and colleagues, working with 758 consultants at Boston Consulting Group, named what they found the jagged technological frontier. Inside it, on tasks the model handled well, AI-assisted consultants were dramatically better and faster. Outside it, on a task designed to sit just beyond the model's competence, consultants using AI did worse than consultants with no AI at all, because they accepted confident output they should have questioned.
Then there is the finding that should trouble anyone planning a training programme. Yu and colleagues, in Nature Medicine in 2024, randomised AI assistance across 140 radiologists and roughly 5,190 observations. The effect of that assistance diverged sharply between individuals, from strongly positive to strongly negative. Experience did not predict who would benefit. Nor did subspecialty. Nor did prior familiarity with AI. Lower performers did not consistently gain, which is the opposite of the story usually told about AI as a leveller. This is 140 radiologists on 15 chest X-ray tasks, and the authors do not claim the pattern holds outside diagnostic imaging.
Put the three together and the position is specific. On average, across a large body of experiments, pairing a human with AI makes decisions worse rather than better. In the field, the sign of the effect flips depending on whether the task sits inside the model's competence. And in the one setting where individual variation has been measured carefully, no available variable predicted which way it would go for a given clinician. None of that argues against using AI. It argues that the design of the relationship carries the outcome, which is the whole of drift versus design.
The clearest measurements of capability loss#
Most of the evidence in this area measures thinking while the tool is present. The harder and more important question is what remains when it is taken away. Two studies now answer that directly.
Budzyń and colleagues, in The Lancet Gastroenterology and Hepatology in 2025, examined 1,443 colonoscopies performed without AI assistance across four Polish centres, 795 before an AI detection tool was introduced and 648 after. The nineteen endoscopists involved averaged twenty-eight years of experience. Their unassisted adenoma detection rate fell from 28.4 per cent to 22.4 per cent, a drop of 6.0 percentage points, with an adjusted odds ratio of 0.69 and a p-value of 0.0089.
This is the second half of the proposition at the top of this page, observed in the field: with the tool absent, what remained was worse than what had been there before it arrived. The study measured unassisted procedures only, so it says nothing about performance while the AI was running, and nothing here should be read as claiming otherwise. The authors are careful and so should anyone quoting them be. It is observational rather than randomised, other changes over the period cannot be fully excluded, it is one procedure in one country, and detection rate is a proxy for skill rather than skill itself. It moves the argument from mechanism to measurement, in a profession where the cost of a degraded practitioner is not abstract.
Bastani and colleagues, in PNAS in 2025, ran the cleaner design. Nearly a thousand high-school students were randomised into three arms: unrestricted GPT-4, a hints-only tutor with guardrails, and a control group with neither. While the tools were present, grades rose 48 per cent with unrestricted access and 127 per cent with the tutor. With access removed, the unrestricted group scored 17 per cent lower than students who had never had it. The guardrailed tutor largely removed that harm.
That last clause is the most actionable finding on this page. It is routinely dropped when the study is cited. The difference between the arms was the interface rather than the model. A tool that gave answers produced a measurable capability deficit. A tool that gave hints, tested on a comparable group in the same experiment, did not. What the study cannot tell you is what a guardrailed interface should look like for professional work, because it tested school mathematics over a bounded period. The principle transfers. The design does not, yet.
A third study measures something adjacent and subtler. Melumad and Yun, in PNAS Nexus in 2025, held the facts constant and varied only the format in which they were delivered. Participants given an AI summary rather than a list of links spent 83.65 seconds engaging with the results against 124.32 seconds, reported learning less and owning the knowledge less, and produced shorter, less specific advice for a friend. The advice they wrote also converged: pairwise similarity between participants rose sharply. Every measure of depth here is self-reported, no experiment tested recall, and the authors describe time on task as a proxy for effort rather than a measure of it. The topics were practical how-to tasks over short horizons, some distance from professional judgement. What is measured directly is time, and time fell by a third.
Where AI helps, and what it is actually helping with#
A page that only collected the losses would be dishonest, and useless to anyone deciding what to deploy. The gains are real and in places very well evidenced. They are also narrower than the enthusiasm around them, so being exact about what has been shown to improve matters here, because in most of these studies it is the output or the system rather than the person's judgement.
The strongest single body of evidence is the MASAI trial in Sweden, which earns attention by being a randomised controlled trial at population scale rather than a laboratory task. The interim safety analysis, covering 80,033 women, found cancer detection of six per 1,000 screened with AI support against five per 1,000 with standard double reading, 41 more cancers found, with an identical false-positive rate of 1.5 per cent in both arms and screen-reading workload down 44 per cent. The full results, published in The Lancet in January 2026 with over 100,000 women and two years of follow-up, found interval cancers falling from 1.76 to 1.55 per 1,000 women, with 16 per cent fewer invasive, 21 per cent fewer large and 27 per cent fewer aggressive-subtype interval cancers.
That is a serious result from a serious design, and the reading it supports is specific. The radiologists were moderately to highly experienced. The AI sat inside a defined workflow with a defined role. What improved is the detection performance of the screening programme. Reading volume fell by 44 per cent, so the humans in it did substantially less reading, and the trial was not designed to test what that does to a radiologist over years. It is also one mammography device, one AI system and one country, which the authors state as limits. The interim analysis did not test patient benefit, mortality was not an endpoint in the full trial and cost-effectiveness was not assessed.
A second result is instructive for the opposite reason. Goh and colleagues, in a randomised clinical trial published in JAMA Network Open in 2024, gave fifty physicians either GPT-4 or conventional resources for diagnostic reasoning. The difference in reasoning scores was two percentage points, with a confidence interval spanning zero. The striking number sits elsewhere in the same trial: the model working alone scored a median 92 per cent, 16 percentage points above the physicians working with conventional resources. The tool outscored the people on this task, and giving it to them changed almost nothing. The authors offer prompting and interaction design as possible explanations rather than established causes, and they explicitly reject the idea that models should diagnose autonomously. The task was six curated vignettes, which exclude history-taking, examination, context and time, and that is most of clinical reasoning.
Order turns out to matter more than most deployments assume. In a study of veterinary radiologists reviewing X-rays, final diagnoses matched the AI 91 per cent of the time when the AI was seen first, against 89 per cent when the clinician committed to a provisional view first, and where the AI flagged a finding the figures were 71 against 65 per cent. The same study found the anchoring produced only marginal diagnostic gains, because of over-reliance on erroneous advice, which is the result that actually bears on whether the workflow helps. It is a working paper, nineteen participants in one speciality, and the effect sizes are small. So this is not a finding that ordering the workflow improves accuracy. What it supports is narrower: forming a view before seeing the machine's answer preserves an independent judgement to compare against, which is worth having whether or not it moves the score. That is the argument developed at human at the start.
The productivity picture is the one most often asserted and least often checked. Brynjolfsson, Li and Raymond, studying 5,172 customer-support agents, found that access to an AI assistant raised productivity by 15 per cent on average, and by 30 per cent for the newest and least experienced staff, while barely moving the most skilled. That is a real gain. The authors are clear it measures output rather than development, over months rather than years, so it does not tell you whether those novices became experts.
Set beside it three pieces of evidence that rarely travel with it, of three different kinds. A randomised trial of 16 experienced open-source developers found them 19 per cent slower when permitted to use AI tools, confidence interval 2 to 39, having forecast a 24 per cent speed-up and still believing afterwards that they had been 20 per cent faster. METR withdrew that figure as a current estimate, and their larger 2026 follow-up points the other way: raw results estimate a speed-up of 18 per cent for returning developers and 4 per cent for new recruits, with every confidence interval crossing zero, and METR themselves believe developers are likely more sped up now than they were a year earlier. A working paper linking adoption surveys to administrative records for roughly 25,000 Danish workers found precise null effects on earnings and hours two years after ChatGPT, ruling out effects larger than 2 per cent, alongside substantial task reorganisation and new work in AI oversight. Denmark is high-trust, high-wage and heavily unionised, and two years is early. And a task-based macroeconomic model, which estimates rather than measures, puts total factor productivity gains at no more than 0.66 per cent over ten years. That model works through task-level cost savings and would not capture effects running through new products, new tasks or capability change.
Those are different settings measuring different things, and none of them refutes the customer-support result. Together they establish something this page needs to be honest about: the productivity gains are real in some settings, absent in others, and considerably less settled than the confident version in circulation. The question this page exists to ask survives either way. If the tool carries a novice to expert-looking output, what happens to the experience through which the novice was supposed to become an expert?
Why it is invisible while it is happening#
If judgement were degrading visibly, this would be a solved problem. Three separate mechanisms conspire to hide it.
The first is metacognitive. Fisher, Goddu and Keil, across nine experiments with 1,708 participants, found that searching the internet inflated people's ratings of their own ability to explain things, at Cohen's d between 0.35 and 0.63, across six unrelated domains. The effect persisted when the search returned no answer to the question asked, and even when it returned no results at all. Access to information was being read as knowledge held internally. Every dependent measure was a self-rating and none tested real knowledge, so the finding is about confidence rather than competence. The authors add that their participants were presumably heavier internet users than average. Confidence is the thing that decides whether anyone checks.
The second is reliability itself, and the sharpest description of it comes from shipping rather than from AI research. A joint safety study by the UK Marine Accident Investigation Branch and its Danish counterpart, published in 2021 after a run of groundings involving electronic chart systems, interviewed 155 deck officers and observed 31 ships at sea. It is one of only two sources on this page not yet carried in the evidence base, so it has not been through the grading this site applies to everything else. Their conclusion is worth quoting exactly: distrust of the instrument, "which is traditionally expected of OOWs, is challenged, because such discrepancies are rarely encountered."
Read that twice, because it inverts the usual worry. The problem was not an unreliable system. The problem was a system reliable enough, often enough, that the habit of checking it stopped being exercised and then stopped existing. Scepticism is a practice rather than a personality trait, and practices decay without occasions to perform them. The officers had been trained. What they lacked was reasons to doubt. Molloy and Parasuraman found the laboratory version of the same thing in 1996: detection of a single automation failure degrades with time on task, and the effect is strongest where the automation has been consistently reliable.
The Dutch Safety Board found something adjacent investigating the grounding of the Nova Cura in the Mytilini Strait, in a report titled, pointedly, Digital navigation: old skills in new technology. Part of the cause was the interface: a crisply rendered display that zoomed smoothly to fine detail, over chart data rated at the lowest reliability category available. The screen looked precise. The underlying survey was not. Confidence had been designed into the presentation rather than earned by the data, which describes a great deal of AI output.
The third mechanism is that the work still ships. Lee and colleagues at Microsoft Research and Carnegie Mellon surveyed 319 knowledge workers about 936 real uses of AI in their jobs, and found that higher confidence in the tool was associated with less critical thinking, and that the thinking which remains shifts in character: from producing to verifying, and from solving to integrating. Critical thinking does not disappear. It moves, and it thins. Nothing in a management system detects that, because every output-based measure reads it as an improvement.
The newer problem: a machine that agrees with you#
The literature above is largely about a human failing to challenge a machine. A second body of evidence, most of it published in the last two years, describes the machine failing to challenge the human, and it changes the shape of the problem.
Sharma and colleagues tested five production AI assistants across four free-form generation tasks and found that all five consistently exhibited sycophancy. The mechanism is in the training rather than the model: both humans and the preference models trained on human judgements prefer convincingly written sycophantic responses over correct ones a non-negligible share of the time, and optimising against those preferences sometimes sacrifices truthfulness. This is a structural consequence of learning from what people like, rather than a defect in one product.
Cheng and colleagues, in Science in 2026, measured the effect on users. Across eleven models, models affirmed users' actions about 50 per cent more often than humans did, including 47 per cent endorsement on prompts describing clearly harmful behaviour. In two preregistered experiments with 1,604 participants, interacting with a sycophantic model reduced willingness to repair an interpersonal conflict and increased conviction of being in the right. Participants rated the sycophantic model as higher quality and trusted it more. The scenarios were interpersonal advice rather than technical judgement, so the transfer to professional decisions is an inference and not a finding.
Two practical notes follow, and their status should travel with them. The first is a company's account of its own incident, independently verified by nobody. OpenAI's own incident report on GPT-4o in 2025 attributed a sycophancy regression to weighting short-term user feedback too heavily, and recorded that offline evaluations and A/B tests looked positive throughout while informal qualitative checks that flagged the problem were overridden. The second is a working paper rather than a peer-reviewed result. Dubois and colleagues found that framing input as a statement rather than a question raised sycophancy by roughly 24 percentage points, and that prompting the model to convert a statement into a question before answering reduced sycophancy more than instructing it not to be sycophantic. They did not test adversarial or devil's advocate prompting, which is the technique most commonly recommended, and no controlled evidence for it was found elsewhere either.
The practical consequence is at how do I get AI to challenge me. The consequence for this page is narrower than it is tempting to make it. Sycophancy is well established as model behaviour, and its measured effect on users so far concerns interpersonal advice rather than technical or analytical decisions. Both authors say so. What follows is a hypothesis this research holds rather than a finding it can cite: a system trained on what people like will tend to reduce the friction that would otherwise force a second thought, and friction is where judgement gets exercised. That should be tested rather than assumed.
Why awareness training is a weak control#
The standard organisational response to everything above is awareness training. The evidence on that is unusually clear and unusually ignored.
Dzindolet and colleagues, in 2003, explained to participants why an automated aid might err. Reliance on it went up rather than down. Explaining the failure modes restored trust even where that trust was unwarranted. The effect is specific, concerning the restoration of trust after an observed error rather than transparency in general. That is still enough to show that telling people about a bias does not remove it.
Parasuraman and Manzey, reviewing decades of work across aviation, medicine and the military, concluded that automation bias and complacency appear in novices and experts alike, resist training, and worsen under workload. That review predates generative AI, and its authors would not claim the effect sizes carry across to a technology this much less predictable. Skitka, Mosier and Burdick had already shown that the failure has a structure: errors of omission, missing what the automation failed to flag, and errors of commission, following automated advice that was wrong. Those two need different countermeasures, and most policies address neither. Skitka's work was a simulated flight task, and whether the effect sizes transfer to generative AI in real professional settings has not been established.
Yu's radiology study does not test training or automation bias, so it cannot be enlisted here as though it did. What it adds is adjacent and still uncomfortable: in the one setting where individual variation has been examined carefully, no available characteristic predicted who would benefit. Taken together, the evidence says you cannot train the bias away, and in at least one clinical setting you cannot identify in advance the people for whom assistance will make things worse. What remains is design: changing the conditions under which the decision is made rather than the disposition of the person making it. Parasuraman and Riley set out the four failure modes in 1997 and the least discussed is still the most relevant here, abuse, meaning automating without regard for the human consequences, which locates the failure with the deploying organisation rather than the individual operator.
Where the evidence remains uncertain#
The direction of the risk is well supported. Its size is not, and being honest about that matters more than a confident headline.
Much of the most-quoted recent work is self-reported or correlational. The Microsoft and Carnegie Mellon study asked workers to describe their own thinking, and people who already think differently may use AI differently; a survey cannot separate the two. Gerlich's 2025 study of 666 participants shows a negative correlation between frequent AI use and critical-thinking scores, mediated by offloading and strongest among the youngest users. It shows a correlation. It also carries a published correction, in Societies in September 2025, which anyone citing it should read alongside it.
The MIT Media Lab EEG study that gave the phrase cognitive debt its currency deserves particular care, because it is the most over-quoted result in the field. It found the weakest brain connectivity and the lowest sense of ownership in the group writing with a language model. It also rests on 54 participants and remains a preprint. A methodological critique from researchers at Vienna and TU Dresden argues it is underpowered, requiring roughly 159 participants against the 54 used, with some figures interpreted from subsamples of two to four essays. The critique is itself an unreviewed preprint offering no competing data, and its sharpest point is conceptual: the search engine group also relied on an external tool and showed no impairment, which complicates any simple offloading story. Treat anyone citing the original as conclusive with the caution its own authors would want. The vocabulary question is mapped at cognitive debt and capability debt.
A 2026 longitudinal pilot found daily AI use rising from 52.4 to 95.7 per cent across three waves, with verification confidence falling and belief-performance gaps widening as tasks got harder. Its authors list its limitations in their own abstract: convenience sampling from a single academic cohort, self-report, no control condition, mathematical problems only, and a timeframe too short for skill trajectories. It is a design precedent for measuring verification rather than a finding about the world.
The productivity evidence, set out above, is genuinely mixed rather than settled, and this page should not be read as leaning on it either way. The capability question runs on a longer clock, across the years over which judgement is actually built or lost, and that is the horizon no workplace study has yet had time to measure. The navigation and memory research is the closest long-run analogue available, and it points one way, but reasoning is not spatial memory and the analogy should carry weight without carrying certainty. The responsible position: the risk is real enough, and slow enough to be invisible, that the time to design against it is now rather than after a decade of proof arrives.
What accumulates when nobody is looking#
The popular fear is that AI will take your job. The more immediate risk is that it takes your judgement, and the two move at very different speeds. Jobs change slowly and visibly, through restructures and headcount, where somebody at least has to decide. Judgement erodes by default, one delegated decision at a time, with nobody choosing it and nobody able to point to the moment it happened. No leadership team sets out to hollow out its own people. It happens because the tool arrives faster than the design.
Underneath that drift are four mechanisms, each developed in its own right elsewhere in this research.
- The missed reps are the repetitions through which judgement was built, now handed to the machine, so the work ships and the practice never happens.
- The missing rungs are the junior tasks that used to carry people up to senior judgement, removed by automation before anyone noticed they were load-bearing.
- Synthetic seniority is the result at the level of the individual: output that looks like judgement without the judgement underneath it.
- Capability debt is the accumulated organisational version: the loss of human knowledge, skill and judgement that builds up when an organisation automates work faster than it redesigns how people learn by doing.
Capability debt is invisible on any dashboard, because the outputs still look fine, right up until a decision arrives that the AI cannot make and no human in the room has been kept capable of making.
What leaders should do#
Decide where human judgement must remain, in advance and in writing. The question is no longer whether to adopt AI. It is which decisions a human has to own, and why. An organisation that has never written that list has already answered by default, one busy afternoon at a time. The working artefact is at the delegation boundary map.
Design the interface, not the policy. This is the finding from Bastani that almost nobody acts on: unrestricted access produced a measurable capability deficit and a hints-only tutor did not. The same model, comparable students, a different interface, and the harm largely gone. Before writing another acceptable-use policy, ask what your tools hand people by default, because that is what the policy is competing with.
Protect the repetitions. If juniors never do the task the AI now does, they never build the judgement the senior role will demand of them. That does not mean banning the tool, which is neither enforceable nor wise. It means keeping deliberate practice in the system on purpose: some work done unaided so the capability is exercised, some AI-assisted work annotated so the person can say what they prompted, what the model returned and what they changed, and some judgement tested directly rather than inferred from a polished output.
Treat verification as real work rather than residue. The jagged-frontier result is the warning to keep in view: the people who trusted AI outside its competence did worse than people with no AI at all. Verification is the judgement layer, it is often harder than production, and in most organisations nobody owns it. The uncomfortable part of that study is that it cannot tell you where the frontier runs in your domain. That is local, it moves with each model release, and it has to be learned rather than looked up. Staff it, train it and value it accordingly, or you will pay least for the work you depend on most. The ownership question is at who owns verification.
Measure capability, not only output. Output quality has stopped being a reliable proxy for the capability of the person who submitted it, which breaks the assumption most promotion and performance systems rest on. Borrow from the professions that solved this long ago: medicine and aviation test judgement directly, through live decisions, simulation and oral examination, rather than trusting that good work implies a capable person. The method is at how do you assess capability rather than output, and the professional precedent at what professions can learn from aviation.
Employers say they want this, which is weaker evidence than it sounds and still worth having. The World Economic Forum's Future of Jobs Report 2025, a survey of employer expectations to 2030, names analytical thinking as the single most valued core skill among employers, and skills gaps as the single biggest barrier to business transformation. It records stated preference rather than hiring behaviour, and the two diverge routinely. The organisations that come out ahead will not be the ones that adopted AI fastest. They will be the ones that decided, deliberately, where human judgement belongs, and built the practice to keep it.
Key research and primary sources
Where a claim matters, go to the study rather than to the article reporting it. Each entry below links to its graded record in the evidence base, with the method, what it supports and what it does not.
- Vaccaro, M., Almaatouq, A. and Malone, T. (2024). When combinations of humans and AI are useful: a systematic review and meta-analysis. Nature Human Behaviour, 8, 2293-2303.
- Budzyń, K., Romańczyk, M., Kitala, D. et al. (2025). Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy. The Lancet Gastroenterology and Hepatology, 10(10), 896-903.
- Bastani, H., Bastani, O., Sungu, A. et al. (2025). Generative AI Without Guardrails Can Harm Learning. PNAS, 122(26).
- Yu, F., Moehring, A., Banerjee, O. et al. (2024). Heterogeneity and predictors of the effects of AI assistance on radiologists. Nature Medicine, 30(3), 837-849.
- Gommers, J., Lang, K., Hofvind, S. et al. (2026). Interval cancer, sensitivity and specificity in the MASAI study. The Lancet, 29 January 2026.
- Lang, K., Josefsson, V., Larsson, A.-M. et al. (2023). MASAI clinical safety analysis. The Lancet Oncology, 24(8), 936-944.
- Goh, E., Gallo, R., Hom, J. et al. (2024). Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial. JAMA Network Open, 7(10).
- Cheng, M., Lee, C., Khadpe, P. et al. (2026). Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence. Science.
- Sharma, M., Tong, M., Korbak, T. et al. (2023). Towards Understanding Sycophancy in Language Models. ICLR 2024.
- Dubois, M., Ududec, C., Summerfield, C. and Luettgau, L. (2026). Ask don't tell: Reducing sycophancy in large language models.
- Fisher, M., Goddu, M. K. and Keil, F. C. (2015). Searching for explanations: How the Internet inflates estimates of internal knowledge. JEP: General, 144(3), 674-687.
- Melumad, S. and Yun, J. H. (2025). Effects of large language models versus web search on depth of learning. PNAS Nexus, 4(10).
- Dell'Acqua, F. et al. (2023). Navigating the Jagged Technological Frontier. Harvard Business School and BCG working paper.
- Brynjolfsson, E., Li, D. and Raymond, L. (2023). Generative AI at Work. NBER Working Paper 31161.
- Model Evaluation and Threat Research (2025). Randomised trial of experienced open-source developers, measured 19 per cent slower with AI tools. Withdrawn by METR as a current estimate.
- Becker, J., Rush, N., Cunningham, T. et al. (2026). The larger follow-up study, in which every confidence interval crosses zero.
- Humlum, A. and Vestergaard, E. (2025). Roughly 25,000 Danish workers across 7,000 workplaces: precise null effects on earnings and hours.
- Acemoglu, D. (2024). The Simple Macroeconomics of AI. Total factor productivity gains of no more than 0.66 per cent over ten years.
- Lee, H.-P. et al. (2025). The Impact of Generative AI on Critical Thinking. Microsoft Research and Carnegie Mellon, CHI 2025.
- Dzindolet, M. T., Peterson, S. A., Pomranky, R. A. et al. (2003). The role of trust in automation reliance. IJHCS, 58(6).
- Parasuraman, R. and Manzey, D. H. (2010). Complacency and Bias in Human Use of Automation. Human Factors, 52(3).
- Skitka, L. J., Mosier, K. L. and Burdick, M. (1999). Does automation bias decision-making? IJHCS, 51(5).
- Molloy, R. and Parasuraman, R. (1996). Monitoring an Automated System for a Single Failure. Human Factors, 38(2).
- Parasuraman, R. and Riley, V. (1997). Humans and Automation: Use, Misuse, Disuse, Abuse. Human Factors, 39(2).
- Risko, E. F. and Gilbert, S. J. (2016). Cognitive Offloading. Trends in Cognitive Sciences, 20(9).
- Sparrow, B., Liu, J. and Wegner, D. M. (2011). Google Effects on Memory. Science, 333(6043).
- Dahmani, L. and Bohbot, V. D. (2020). Habitual use of GPS negatively impacts spatial memory. Scientific Reports, 10, 6310.
- Gerlich, M. (2025). AI Tools in Society. Societies, 15(1), 6. Read with the correction published in September 2025.
- Kosmyna, N. et al. (2025). Your Brain on ChatGPT. MIT Media Lab preprint, and the methodological critique that should be read with it.
- Who Goes First? Influences of Human-AI Workflow on Decision Making in Clinical Imaging (2022).
- Huemmer, M., Durner, F., Shyiramunda, T. and Cummings-Koether, M. J. (2026). AI, Metacognition, and the Verification Bottleneck.
- World Economic Forum (2025). The Future of Jobs Report 2025.
- Marine Accident Investigation Branch (UK) and Danish Maritime Accident Investigation Board (2021). Application and Usability of ECDIS: a joint safety study. Not yet carried in the evidence base.
- Dutch Safety Board. Digital navigation: old skills in new technology, on the grounding of the Nova Cura. Not yet carried in the evidence base.
Related SuperSkills research#
On the choice that decides the outcome, drift versus design. On how judgement is built rather than lost, how humans learn with AI and using AI without dependency. On splitting decisions between people and machines, human and AI decision making and decision quality. On what remains distinctly human, what stays human. The underlying mechanism is defined at cognitive offloading, the failure mode at automation bias, and the recognition question at outsourced recognition. On spotting error in practice, how do I know when AI is wrong. On whether oversight is real, human in the loop is not a safeguard. Every study behind this page, with its method and its limits, is in the evidence base; the claims themselves, banded by strength and including what remains unknown, are at what we actually know.
About this research#
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. This work draws on research across more than 200 organisations in 30 countries over seven years. The page is written to a deliberate rule: findings are attributed to the studies that produced them and kept separate from the interpretation, which is the author's. The named concepts, drift versus design, the missed reps, the missing rungs, synthetic seniority and capability debt, are part of the SuperSkills lexicon. This is a living reference, reviewed and updated as significant new evidence appears, rather than a dated article left to stand.
Cite this
Hirji, R. (2026). Does AI weaken human judgement? The evidence, checked. The SuperSkills Intelligence Company. Last reviewed 29 August 2026. thesuperskills.com/research/ai-and-human-judgement
