← Research
Research · Evidence

The evidence on AI and human capability

A curated reference to the best research in the field: what each study found, what it shows, and what it does not prove.

Last reviewed: 26 August 2026

This is a living reference to the strongest research on how AI is changing human capability, curated and interpreted rather than simply listed. Each study is given with its finding, what it establishes, what it does not, and a link to the original. It is the evidence base beneath the rest of this research.

For a decade this field has been read here every week, sifting the serious research from the noise. This page is the result: a single, honest place that references the best work on AI and human capability, from the institutions leaders already trust and the academic studies that underpin them. The rule throughout is simple. The evidence is kept separate from the interpretation. Each entry states what a study actually found, then what that does and does not license anyone to conclude, and links straight to the source so you can check it yourself.

How to read this

Every entry follows the same shape: the finding, then what it shows, then what it does not prove. That second line matters most. Most commentary on AI and human capability fails by treating a suggestive result as a settled fact, in either direction. Kept honest, the studies below are more useful and more persuasive than any single headline drawn from them. This is a living document; studies are added as they are verified, and a weekly monitor flags new work from the major research bodies for inclusion.

Judgement and critical thinking

Lee, Sarkar and colleagues (2025), Microsoft Research and Carnegie Mellon. In a survey of 319 knowledge workers, higher confidence in the AI was associated with less critical thinking, while higher confidence in one's own ability was associated with more of it. What it shows: in real workplace use, reliance and critical engagement trade off against each other. What it does not prove: it is self-reported and correlational, a snapshot rather than a measure of skill changing over time. Microsoft Research

Gerlich (2025), Societies. Across 666 participants, frequent AI-tool use correlated with weaker critical-thinking scores, an effect statistically mediated by cognitive offloading and strongest among younger users. What it shows: a plausible mechanism linking habitual offloading to reduced critical thinking. What it does not prove: causation or its direction; heavy offloaders may already differ from light ones. Societies 15(1), 6

Kosmyna and colleagues (2025), MIT Media Lab. Using EEG across repeated essay-writing sessions, people who wrote with a large language model showed lower neural connectivity, remembered less of what they had written, and felt less ownership of it than those who wrote unaided. What it shows: measurable, in-the-moment differences in engagement when the thinking is offloaded. What it does not prove: lasting cognitive harm; the sample is small and the horizon short. MIT Media Lab

Cognitive offloading and memory

Sparrow, Liu and Wegner (2011), Science. People who expected to be able to look information up later remembered the information itself less well, but remembered where to find it, the effect now known as the Google effect. What it shows: external availability changes what the mind bothers to retain. What it does not prove: that this is harmful rather than an efficient reallocation of limited memory. Science

Risko and Gilbert (2016), Trends in Cognitive Sciences. A review that defines cognitive offloading, the use of external tools to reduce mental effort, and sets out when people choose to offload and what it costs them. What it shows: a rigorous framework for the entire question of leaning on external aids. What it does not prove: anything specific to AI; it predates today's tools, but explains them well. Trends in Cognitive Sciences

Dahmani and Bohbot (2020), Scientific Reports. In adults, greater lifetime GPS use was associated with worse spatial memory, and heavier users declined faster over the course of the study. What it shows: a real-world case in which habitual reliance on a tool tracks a specific human capability falling. What it does not prove: that AI will affect judgement the same way; navigation is one narrow skill. Scientific Reports

Expertise, novices and the jagged frontier

Dell'Acqua and colleagues (2023), Harvard Business School and BCG. Among 758 consultants, those using GPT-4 on tasks inside its competence were markedly more productive and higher in quality, but on a task designed to fall just outside it, AI users were wrong more often than those working with no AI at all. What it shows: AI helps inside a jagged frontier of competence and quietly misleads just beyond it. What it does not prove: where that frontier sits for any given real-world job. Harvard and BCG

Brynjolfsson, Li and Raymond (2023), NBER. Among 5,179 customer-support agents, access to a generative-AI assistant raised productivity by about 14 percent on average and 34 percent for the least experienced, with little effect on the most skilled. What it shows: AI can transfer expert patterns to novices, a real and immediate gain. What it does not prove: that novices retain the underlying judgement once the tool is taken away. NBER w31161

Automation, oversight and accountability

Parasuraman and Manzey (2010), Human Factors. A review of decades of research on automation bias and complacency: people tend to under-monitor reliable automation and over-trust its output, especially under load. What it shows: over-reliance is a long-established human tendency, not a new problem invented by AI. What it does not prove: the size of the effect for modern AI, which is far more capable and persuasive than the systems studied. Human Factors

The SCHUFA judgment (2023), Court of Justice of the European Union, C-634/21. The Court held that an automated credit score a lender relies on decisively can itself amount to a decision under the GDPR, bringing it within the rules on automated decision-making. What it shows: accountability can attach to a model's output, not only to the human who signs it off. What it does not prove: how the principle resolves in practice across different sectors and systems. CURIA

The EU AI Act (2024), Regulation 2024/1689. The first comprehensive AI law requires, for high-risk systems, effective human oversight by people able to understand, monitor and override the system. What it shows: that meaningful human oversight is becoming a legal requirement, not merely good practice. What it does not prove: how regulators will judge oversight to be effective rather than a rubber stamp. EUR-Lex

Jobs, skills and the workforce

World Economic Forum (2025), Future of Jobs Report. Employers name skills gaps as the single biggest barrier to business transformation over the next five years, and rank analytical thinking, resilience, flexibility and curiosity among the most important skills. What it shows: the reported bottleneck is human capability, not the tooling. What it does not prove: outcomes; it is a large survey of employer expectation rather than an audit of what happened. World Economic Forum

PwC (2025), Global AI Jobs Barometer. Analysing close to a billion job advertisements, PwC found that industries most exposed to AI showed roughly three times higher growth in revenue per employee, carried wage premiums for AI-skilled workers, and demanded skills that are changing about 66 percent faster than elsewhere. What it shows: AI is reshaping the value and content of work, not simply removing it. What it does not prove: how those gains are distributed, or the effect on entry-level pipelines. PwC

EY (2025), Work Reimagined. Across 15,000 employees and 1,500 employers in 29 countries, 88 percent use AI at work but only 5 percent in genuinely transformative ways, 37 percent fear over-reliance is eroding their skills, and organisations on weak talent foundations forfeited up to 40 percent of the productivity available to them. What it shows: the human foundation, not the tool, gates the return on AI. What it does not prove: a direct measure of skill loss, which remains unbuilt. EY

What the weight of evidence suggests

Read together, and read honestly, these studies tell a coherent story. AI reliably improves performance on the tasks it suits, and it helps the least experienced most, which is a genuine and immediate gain. The difficulty is that the risks sit in exactly the same places as the gains. Performance degrades just beyond the model's competence, where people trust output they cannot check. Habitual offloading tracks weaker engagement, memory and critical thinking in the moment. And the reported constraint on organisations is not the technology but the capability of the people using it. No single study here proves long-run deskilling, and it would be dishonest to claim one does. What the body of evidence supports is narrower and more useful: capability erodes when the doing is automated without redesigning the learning, and that outcome is a choice rather than a certainty.

How this connects to the SuperSkills research

This canon is the evidence base beneath the rest of the work. The trade-off between reliance and critical thinking runs through AI and human judgement and AI and critical thinking. The novice-to-expert transfer and the practice it skips are the mechanism behind synthetic seniority, the missing rungs and capability debt. The oversight and accountability line sits under human at the start and AI agents and human judgement. And the whole reading, that this is avoidable by design, is the argument of drift versus design.

About this reference

Rahim Hirji is the author of SuperSkills: The Seven Human Skills for the Age of AI (Kogan Page, 2026) and the founder of The SuperSkills Intelligence Company. This reference is curated in the spirit of a decade of weekly reading of this field. Findings are attributed to the studies that produced them and kept separate from the interpretation, which is the author's. Links go to the original source in every case. This is a living document, reviewed and extended as significant new evidence is verified.

Cite this

Hirji, R. (2026). The evidence on AI and human capability. The SuperSkills Intelligence Company. Last reviewed 26 August 2026. thesuperskills.com/research/evidence

In this hub

The Evidence Base

The research base beneath the work. It underpins these hubs:

The work

Where the writing comes from.

These essays draw on research across more than 200 organisations in 30 countries. See the wider body of work, or bring it into your organisation.

All research →