The most useful thing anyone can do in a noisy field is separate what we know from what we suspect from what we do not know yet. Almost nobody does it, because the first category is smaller than the discourse requires and the third is larger than anyone selling a solution would like. This page is that separation, run against the 56 graded studies in the evidence base. Nineteen claims, three bands, every one linked to its source and labelled with the strength of evidence behind it. Reviewed quarterly, with changes logged rather than quietly made.
How claims are graded
Tier A peer-reviewed research, randomised trials, systematic reviews and meta-analyses, official statistics. Tier B credible working papers, large field experiments, institutional research with a transparent method. Tier C corporate and institutional surveys, commercial labour-market data, practitioner research. Tier D expert interpretation, books, journalism, individual cases.
A claim resting on Tier D is not written with the confidence of a claim resting on several Tier A studies. Strong below means multiple independent sources pointing the same way with no serious contrary finding. Emerging means real findings that are single, narrow or unreplicated. Unknown means the question is being answered confidently in public and the evidence does not answer it.
Strong evidence
Seven claims. Each rests on multiple independent sources, mostly Tier A, pointing in a consistent direction.
1. AI can materially improve performance on some cognitive tasks.
Consultants inside the model's competence completed 12.2 per cent more tasks, 25.1 per cent faster, at more than 40 per cent higher quality (Dell'Acqua 2023). Customer-support productivity rose 14 per cent on average (Brynjolfsson 2023). Mathematics grades rose while the tutor was available (Bastani 2025). This is the least contested claim on the page and it is worth stating plainly, because the rest of this site is often read as scepticism about capability. It is not.
2. The gains skew towards the less experienced, in several independent settings.
34 per cent for the newest support agents against almost nothing for the most skilled (Brynjolfsson 2023); 43 per cent for below-average consultants against 17 per cent for above-average (Dell'Acqua 2023); gains accruing almost entirely to low-skilled drivers, narrowing the spread by 14 per cent (Kanazawa 2022); largest creative gains for the least creative writers (Doshi and Hauser 2024). Four settings, four methods, one direction.
3. The boundary of AI competence is jagged, and invisible from the output.
On a task placed just outside the frontier, AI-assisted consultants were 19 percentage points less likely to reach a correct solution than consultants with no AI at all (Dell'Acqua 2023). The effect of assistance on radiologists ran from strongly positive to strongly negative between individuals and was not predicted by experience or prior familiarity (Yu 2024). Graded B rather than A because the frontier finding sits in a working paper, though a heavily scrutinised one.
4. Automation bias is real, appears in experts as well as novices, and resists training.
A review of two decades of evidence finds complacency and bias in both groups, worsening under workload and not removed by awareness (Parasuraman and Manzey 2010). People often weight algorithmic advice more heavily than human advice, with domain experts the notable exception (Logg 2019). See automation bias.
5. Cognitive offloading is a long-documented human behaviour that predates AI.
Defined and measured well before generative models (Risko and Gilbert 2016); people remember where to find information rather than the information itself when they expect it to remain available (Sparrow 2011); habitual satnav users showed worse unaided spatial memory, with steeper decline over three years of heavier use (Dahmani and Bohbot 2020). Anyone presenting offloading as a novel consequence of AI is not describing the literature. See cognitive offloading.
6. Pairing a human with an AI system does not reliably beat the better of the two alone.
A meta-analysis of 370 effect sizes across 106 experiments found combinations performed worse on average than the stronger party alone, with losses concentrated in decision tasks and gains in creation tasks (Vaccaro 2024). Reinforced by the individual-level divergence in Yu 2024. This is the single most under-known finding in the field, and it undermines the most common deployment pattern. See human-AI collaboration.
7. Measured aggregate labour-market effects so far are far smaller than the public discussion implies.
Precise null effects on earnings and hours two years after ChatGPT, ruling out effects larger than 2 per cent, alongside real task reorganisation (Humlum and Vestergaard 2025). Firm-level adoption remains low in several major economies: 12.9 per cent reporting any firm use in Japan (JILPT 2025), 2.7 per cent of Korean firms of ten or more staff (KDI 2023). Individual use is meanwhile widespread (Bick 2024). The gap between personal adoption and institutional adoption is itself the story. This claim is about what has been measured to date and is not a forecast.
Emerging evidence
Six claims. Real findings, but single studies, narrow tasks, correlational designs, or not yet replicated. These are the ones most likely to change.
8. Heavy reliance may reduce cognitive effort and shift thinking from producing to verifying.
Higher confidence in the tool was associated with less critical thinking, with the remaining work shifting from solving towards integrating (Lee 2025, self-reported survey of 319 knowledge workers). A negative correlation between frequent use and critical-thinking scores, mediated by offloading (Gerlich 2025, correlational, cannot establish direction). Weakest brain connectivity and lowest sense of ownership in the LLM group (Kosmyna 2025, small sample, working paper). Three studies, three weaknesses: self-report, correlation, and sample size. They agree, which is suggestive, but agreement between three weak designs is not strength.
9. Dependency may persist after the tool is removed.
When access was withdrawn, students who had used an unrestricted interface scored 17 per cent lower than students who never had access, while a guardrailed version largely removed the harm (Bastani 2025). Tier A method, but a single study, one subject, one age group. It is the most important unreplicated finding in the field and the one most worth watching.
10. AI may reduce collective diversity even while raising individual quality.
293 writers and 600 evaluators: AI-assisted stories were rated better and were markedly more similar to one another (Doshi and Hauser 2024). Strong design, one creative domain. Whether this generalises to engineering, law or strategy is untested. See does AI make everyone think alike?
11. Routine AI assistance may erode unassisted clinical skill.
Adenoma detection in unassisted colonoscopy fell from 28.4 per cent before AI exposure to 22.4 per cent after, a drop of 6.0 percentage points (Budzyn 2025). This is the first real-world clinical evidence of the effect and it carries a patient outcome, which makes it the strongest single finding on this site. It is also retrospective and observational across four Polish centres rather than randomised, so change over time from other causes cannot be excluded. A consequential result on a design that cannot yet prove causation.
12. AI may be disrupting the routes through which novices acquire expertise.
Historic robot adoption in Germany left incumbents in higher-quality tasks while the cost fell on young labour-market entrants (Dauth 2021). French employment of under-30s fell 7.4 per cent year on year in IT services (INSEE 2026). Indian graduate unemployment sits near 40 per cent for the under-25s (Azim Premji 2026). Substitutability rose about ten percentage points for degree-level expert occupations while staying flat for helper occupations (IAB 2024). Important caveat: these are exposure and coincidence, not causation. Youth employment is sensitive to interest rates, hiring freezes and cohort size. This is the claim on the page where the interpretation runs furthest ahead of the evidence, including in my own writing. See missing rungs.
13. Which tasks get automated, expert or non-expert, may determine whether wages rise or fall.
Automation that removed less-expert tasks raised wages and reduced employment; automation that removed expert tasks lowered wages and raised employment (Autor and Thompson 2025). A single framework paper with historical support rather than a body of replication, but it reframes the whole "how many jobs" argument into a better question: which tasks, and whose expertise.
Unknown
Six questions being answered confidently in public that the evidence does not answer. Listing them is not modesty. It is the part that makes the rest of the page trustworthy.
14. Whether long-term generative-AI use produces durable deterioration in critical thinking.
No longitudinal study exists. Everything available is cross-sectional, self-reported, short-horizon, or all three. Anyone stating this as established, in either direction, is going beyond the evidence. What would settle it: a multi-year cohort study with objective measures and a plausible control group.
15. Whether AI increases or reduces creativity at population level.
Claim 10 shows convergence in one domain over one task. Extrapolating from that to civilisational creative decline is not supported, and neither is the opposite claim about a creative renaissance. What would settle it: replication across domains, plus a measure of output diversity at field level over several years.
16. Whether capability debt is measurable, or currently a useful metaphor without an instrument.
This is my own concept and it is on this list for that reason. The mechanism is plausible and the component findings exist, but no validated instrument measures organisational capability erosion over time. What would settle it: a repeatable measure of unassisted capability, tracked in the same organisation across several years. Until that exists, capability debt is a framework, not a finding, and this site says so.
17. Whether removing junior work actually impairs the formation of senior capability.
The deliberate-practice literature (Ericsson 1993) supports the mechanism, and the meta-analytic challenge to it (Macnamara and Maitra 2019) weakens how much practice explains. No study has tested whether people deprived of junior repetitions become worse seniors, or whether they compensate through other routes. This is the load-bearing assumption underneath a large part of the argument on this site, and it is currently unproven.
18. Whether human skills are becoming more economically valuable, or only more discussed.
Employer surveys have named analytical and creative thinking as leading skills for eight years running (WEF 2018, 2020, 2023, 2025). That is stated demand, not revealed price. The wage-premium evidence that does exist attaches to AI skills rather than human ones (PwC 2025). What would settle it: wage and vacancy data showing a rising return to these capabilities specifically, rather than employers saying they matter.
19. Whether the frontier smooths as models improve.
The jagged-frontier result used GPT-4 in 2023. Tasks outside the boundary then may sit inside it now. Whether the boundary becomes smoother and more predictable, or simply moves while staying jagged, is unstudied. It matters enormously for whether claim 3 remains true. What would settle it: a repeat of the Dell'Acqua design on current systems.
What would change these positions
Every claim in the emerging band is one good study away from moving in either direction. The ones I am actively watching, and would revise fast on, are: any replication or failure to replicate the Bastani persistence finding; any extension of Doshi and Hauser outside creative writing; any longitudinal work on critical thinking; any methodological challenge to Budzyn that holds; and any repeat of the jagged-frontier experiment on current models.
If a finding here is overturned, it will be corrected on this page with the change dated and the previous position left visible. A page that only ever gains confidence is not tracking evidence, it is tracking commitment.
Related SuperSkills research
The full graded catalogue, with what each study does not support, is in the evidence base. The works that shaped the field are in the canon. Dated positions and how they have held up are in predictions and the timeline. The overall argument is in the SuperSkills thesis. The full map of the territory, including the questions this research has not yet answered, is in 500 Questions About Humans and AI.
About this page
Rahim Hirji is the author of SuperSkills: The Seven Human Skills for the Age of AI (Kogan Page, 2026) and the founder of The SuperSkills Intelligence Company. This page is reviewed quarterly. Claims are graded by evidence strength rather than by how well they support the argument made elsewhere on this site, which is why three of the six unknowns undercut positions taken here. Findings are attributed to the studies that produced them and kept separate from the interpretation.
Cite this
Hirji, R. (2026). What we actually know about AI and human capability. The SuperSkills Intelligence Company. Last reviewed 26 August 2026. thesuperskills.com/research/what-we-know-about-ai-and-human-capability