- Where is the evidence on AI and human capability?
- What does the evidence on AI and human capability not show?
This is the evidence base behind the SuperSkills research: the studies and reports that bear on what increasingly capable AI does to human capability, each with its method, its finding, and, most importantly, what it does not support. That last field is the reason this exists. Almost every claim you will read about AI and human thinking rests on a handful of papers, and most of them are quoted well beyond what their design can carry. Every entry here has a permanent link, so it can be cited on its own. Every external source has been fetched and confirmed to resolve.
276 entries across 10 sections. This is a long reference and it is meant to be. If you came for one study, use your browser's find, or jump to its section.
- Judgement, thinking and offloading 31
- Learning, practice and expertise 48
- Human and AI collaboration 22
- Empathy, creativity and what stays human 8
- Work, jobs and the labour market 31
- The seven human capabilities 21
- The institutional record 57
- Professions and sectors 14
- Frontline, clinical and physical work 22
- The international picture 22
Every entry has a permanent anchor, so any single study can be linked and cited on its own. The whole base is also machine-readable at evidence.json.
How to read this#
Entries are graded by what kind of evidence they are, because the distinction is routinely collapsed and it matters. A preregistered meta-analysis in Nature Human Behaviour and a consultancy's survey of its own client base are both cited as "research" in the same sentence every day, and they cannot carry the same weight. The grades are deliberately plain.
- Peer-reviewed. Published in a peer-reviewed journal or archival conference.
- Working paper. Circulated for comment; not yet peer-reviewed.
- Institutional survey. Survey by an organisation, usually self-selected respondents, not peer-reviewed.
- Institutional modelling. Projection or secondary analysis by an organisation, assumption-driven.
- Compiled review. Aggregation of third-party data rather than original research.
- Statutory investigation. Investigation of a single event by a public body with access to primary records. Determinations rather than estimates, and N is one.
- Argued perspective. A framework or risk model argued in a reviewed venue, reporting no new data. Carries the authority of its reasoning, not of a measurement.
- Qualitative field study. Interviews, observation or participatory work in a real setting, peer-reviewed. Rich on mechanism and meaning; no measured outcome and no control.
- Vendor research. Empirical work published by a company about the effects of its own product, using proprietary usage data as an input. Often careful and usually unreproducible: the key variable cannot be rebuilt or checked from outside the firm.
- Operator account. An organisation's own account of a system it runs. Capability as claimed rather than measured, with no independent verification and an interest in the result.
Institutional reports are included rather than excluded. They are often the best available data on scale, sentiment and demand, and they are what boards actually read. But several are published by organisations that sell the remedies they recommend, and where that is true it is stated in the entry. No source here is dismissed for its origin, and none is granted authority for it either.
This base is not exhaustive and does not pretend to be. It is the material this research actually rests on, which is a more useful thing than a bibliography.
The source hierarchy#
The grades above describe what a source is. This describes how much weight it is allowed to carry. Both are published so that any claim on this site can be challenged against a stated standard rather than against a pile of links.
Four tiers#
Tier A. Peer-reviewed research, randomised controlled trials, systematic reviews and meta-analyses, official statistics, government and international datasets.
Tier B. Credible working papers, large-scale field experiments, institutional research with a transparent and reproducible method.
Tier C. Corporate and institutional surveys, commercial labour-market datasets, practitioner research, industry bodies.
Tier D. Expert interpretation, books, commentary, journalism, individual cases.
The rule that follows is the important part. A claim resting on Tier D is not written with the confidence of a claim resting on several Tier A studies, and the language on every page tracks the tier: the evidence shows for A, findings suggest for B, reported or self-reported for C, argued or observed for D. Where a page rests on a single study, it says so, and says what would strengthen the claim.
Tier D is not a lesser category to be avoided. Cases, books and reporting are how the international and sectoral research gets its texture, and a well-sourced case is often more useful than a weak survey. It simply cannot carry a load it was not built for. Mapping to the grades above: peer-reviewed is Tier A, working papers and institutional modelling are Tier B, institutional surveys are Tier C, and compiled reviews are C or D depending on method.
For the claims themselves, banded by how much of this base supports each one, see what we actually know about AI and human capability. The full map of the territory, including the questions this research has not yet answered, is in 500 Questions About Humans and AI. The full method, including how these tiers are applied and where AI is used in producing this research, is in how this research works, and every correction to date is logged in corrections.
Judgement, thinking and offloading
What happens to reasoning when a machine will do it for you.
Peer-reviewed#
Mackworth, N.H. (1948). The Breakdown of Vigilance during Prolonged Visual Search#
Quarterly Journal of Experimental Psychology, 1(1), 6-21. DOI 10.1080/17470214808416738. Much of the same material was reported at greater length in Researches on the Measurement of Human Performance, MRC Special Report 268, HMSO, 1950, which is where the 1950 date in circulation comes from
- Method
- The Clock Test, built to simulate radar and sonar watchkeeping. An unmarked clock face whose pointer jumps in equal steps about once a second, making a rare double jump at irregular intervals, which the observer reports by pressing a button. Two-hour watches, analysed in half-hour blocks. Signal probability in the operational task was a little over half a per cent.
- Finding
- Detection of rare signals fell measurably between the first and second half-hour block and continued to decline across the watch. The vigilance decrement, and the founding result of the field.
- What it supports
- That sustained attention to rare events degrades within the first hour, as a property of the task rather than of the person, which is the mechanism behind every oversight arrangement that asks someone to watch a mostly correct system.
- What it does not support
- Anything more precise about timing than the block structure allows. Because Mackworth analysed in half-hour blocks, the decline cannot be located inside the first thirty minutes, and the widely repeated claim that accuracy falls within the first half hour states the result more sharply than the design supports. The specific percentage figures in circulation come from secondary literature: the original is paywalled and was not read at source for this entry.
Peer-reviewed#
Staw, B. M. (1976). Knee-Deep in the Big Muddy: A Study of Escalating Commitment to a Chosen Course of Action#
Organizational Behavior and Human Performance, 16(1), 27-44
- Method
- Role-played corporate financial decision, 240 business students, two-by-two factorial crossing personal responsibility for the initial investment against positive or negative decision consequences.
- Finding
- Participants personally responsible for the earlier investment allocated an average of 11.08 million dollars to the division they had chosen, against 8.89 million where another officer had chosen it. Negative consequences drew 11.20 million against 8.77 million for positive ones. Where a participant's own earlier choice had subsequently declined, the figure rose to 13.07 million. Both main effects and the interaction were significant.
- What it supports
- That responsibility for the original decision changes the next one, in a known direction, before any question of competence arises. This is why asking a programme's sponsor whether to continue it is not an assessment.
- What it does not support
- Anything about real organisations or real money. A single-session paper exercise with undergraduates, no longitudinal element, and no prevalence claim. Staw himself flags the ambiguity between self-justification and self-perception as the mechanism.
Peer-reviewed#
Molloy, R. and Parasuraman, R. (1996). Monitoring an Automated System for a Single Failure: Vigilance and Task Complexity Effects#
Human Factors, 38(2)
- Method
- Laboratory flight-simulation experiments varying task complexity and time on task.
- Finding
- Detection of a single automation failure degrades with time on task, and the effect is strongest where the automation has been consistently reliable.
- What it supports
- That monitoring performance falls predictably rather than randomly, and that reliability itself is part of the cause.
- What it does not support
- Field incidence. It is simulation, not accident data.
Peer-reviewed#
Parasuraman, R. and Riley, V. (1997). Humans and Automation: Use, Misuse, Disuse, Abuse#
Human Factors, 39(2)
- Method
- Review and framework paper synthesising the human-factors literature on how people interact with automated systems.
- Finding
- Establishes four distinct failure modes: use, misuse through over-reliance, disuse through under-reliance, and abuse through automating without regard for the human consequences.
- What it supports
- That over-reliance and under-reliance are separate problems requiring separate design responses, and that the failure can sit with the deploying organisation rather than the operator.
- What it does not support
- Quantified effect sizes. It is a framework paper rather than an experiment.
Peer-reviewed#
Skitka, L. J., Mosier, K. L. and Burdick, M. (1999). Does automation bias decision-making?#
International Journal of Human-Computer Studies, 51(5)
- Method
- Controlled experiments with a simulated flight task, comparing automated and non-automated decision aids.
- Finding
- Automated aids produced two distinct error types: errors of omission, missing events the automation failed to flag, and errors of commission, following automated advice that was wrong.
- What it supports
- That automation bias has a measurable structure, and that the two error types require different countermeasures.
- What it does not support
- That the effect sizes transfer to generative AI or to non-simulated professional settings.
Peer-reviewed#
Dzindolet, M. T., Peterson, S. A., Pomranky, R. A., Pierce, L. G. and Beck, H. P. (2003). The role of trust in automation reliance#
International Journal of Human-Computer Studies, 58(6)
- Method
- Experiments manipulating what participants were told about an automated aid's reliability and failure modes.
- Finding
- Explaining why an automated aid might err INCREASED reliance on it, restoring trust even where that trust was unwarranted.
- What it supports
- That awareness training is a weak control, and can move reliance in the opposite direction to the one intended.
- What it does not support
- That explanation is always counterproductive. The effect is about restoring trust after observed error, not about all forms of transparency.
Peer-reviewed#
Kahneman, D. and Klein, G. (2009). Conditions for intuitive expertise: a failure to disagree#
American Psychologist, 64(6), 515-526
- Method
- Adversarial collaboration between the leading proponents of the heuristics-and-biases and naturalistic-decision-making traditions, who had reached opposite conclusions about expert intuition. A joint theoretical paper rather than a new experiment.
- Finding
- Judging the likely quality of an intuitive judgement requires assessing two things: the predictability of the environment in which the judgement is made, and the individual's opportunity to learn that environment's regularities. Where both hold, recognitional expertise is trustworthy. Where either fails, confident intuition is not evidence of skill.
- What it supports
- That expert intuition is neither reliably good nor reliably poor, and that the discriminating variable is the environment rather than the expert. It also supplies the test for when preserving human judgement is worth the cost.
- What it does not support
- Which specific professional environments meet the conditions. The authors give the criteria and not a classification, so applying it to law, medicine, consulting or management is a judgement in itself.
Peer-reviewed#
Leroy, S. (2009). Why Is It So Hard to Do My Work? The Challenge of Attention Residue When Switching Between Work Tasks#
Organizational Behavior and Human Decision Processes, 109(2)
- Method
- Two laboratory experiments manipulating whether a task was completed or interrupted before switching.
- Finding
- People struggle to move attention away from an unfinished task, and performance on the next task suffers. Time pressure on the first task helps disengagement.
- What it supports
- That the residue effect exists under controlled conditions, and that completion rather than willpower is what releases attention.
- What it does not support
- Real-world magnitude. Two laboratory experiments, not a field study.
Peer-reviewed#
Parasuraman, R. and Manzey, D. H. (2010). Complacency and Bias in Human Use of Automation: An Attentional Integration#
Human Factors, 52(3)
- Method
- Review across aviation, medicine and military domains.
- Finding
- Automation bias and complacency appear in novices and experts alike, resist training, and worsen under workload.
- What it supports
- That under-questioning automated advice is a robust, decades-old finding, not a novelty of the AI era.
- What it does not support
- The size of the effect for generative AI, which is far less predictable than the automation studied here.
Peer-reviewed#
Sparrow, B., Liu, J. and Wegner, D. M. (2011). Google Effects on Memory: Cognitive Consequences of Having Information at Our Fingertips#
Science, 333(6043)
- Method
- Four laboratory experiments.
- Finding
- When people expect information to remain available, they remember where to find it rather than the thing itself.
- What it supports
- That expected availability changes what gets encoded.
- What it does not support
- That total memory capability declines, or that the trade is net negative.
Peer-reviewed#
Budescu, D. V., Por, H.-H. and Broomell, S. B. (2012). Effective communication of uncertainty in the IPCC reports#
Climatic Change, 113, 181-200
- Method
- Nationally representative US survey experiment via TESS and the Knowledge Networks panel, December 2009 to January 2010. 841 invited, 556 completed, 66 per cent. Three between-subjects conditions: control 193, translation table supplied 175, verbal terms with numerical ranges in the text 188. Eight sentences from IPCC reports, four probability terms tested.
- Finding
- Mean estimates were 41 for very unlikely, 44 for unlikely, 54 for likely and 62 for very likely, against IPCC guidelines of below 10, below 33, above 66 and above 90. Consistency with the guidelines was 20.76 per cent in the control, 18.81 per cent when the translation table was supplied and 30.12 per cent when numerical ranges appeared alongside the words. 24 per cent of respondents gave no response consistent with the guidelines and only 6 per cent gave six or more. The authors describe the pattern as regressive, with the median respondent reading an intended 0.90 as about 0.65 to 0.75.
- What it supports
- That publishing a glossary does not fix a probability vocabulary, and that only putting the number in the sentence helps, and then only to under a third consistency.
- What it does not support
- That this settles the design. Only four of the seven IPCC terms were tested, no lower or upper bound data were collected, and the authors decline to read the result as a criticism of the IPCC, noting there is no optimal method. The widely cited 2009 Psychological Science paper by the same authors could not be opened during the 4 September 2026 build, so its figures appear nowhere on this estate.
Peer-reviewed#
Budescu, D. V., Por, H.-H., Broomell, S. B. and Smithson, M. (2014). The interpretation of IPCC probabilistic statements around the world#
Nature Climate Change, 4(6), 508-512
- Method
- Multi-national survey experiment, 25 samples across 24 countries and 17 languages, testing four target terms under a translation condition and a verbal-numerical condition.
- Finding
- Laypeople interpret IPCC statements as conveying probabilities closer to 50 per cent than intended by the IPCC authors. Supplementing verbal terms with numerical ranges increases correspondence with the guidelines and improves differentiation between terms. The authors describe the qualitative patterns as remarkably stable across all samples and languages, and note that interpretations across languages become more similar under the numerical format.
- What it supports
- That the regressive reading of probability words is not an artefact of English or of American respondents. It is general.
- What it does not support
- Any magnitude quotable from this estate. The article is paywalled and only the abstract was read on 4 September 2026, so no participant count, per-country sample or consistency percentage is given anywhere here.
Peer-reviewed#
Fisher, M., Goddu, M. K. and Keil, F. C. (2015). Searching for explanations: How the Internet inflates estimates of internal knowledge#
Journal of Experimental Psychology: General, 144(3), 674-687. DOI 10.1037/xge0000070
- Method
- Nine between-subjects experiments, 1,708 US participants via Amazon Mechanical Turk. An induction phase in which participants either searched the internet for explanations or were told not to, followed by self-ratings of their ability to explain questions in six domains unrelated to the induction material.
- Finding
- Searching inflated self-rated explanatory ability, with Cohen's d from 0.35 to 0.63 across studies, and the effect appeared across all six unrelated domains. It persisted when the search returned no answer to the question asked (Experiment 4b: 4.11 against 4.00 for those who found an answer, both far above a 3.05 no-search baseline) and when it returned no results at all (Experiment 4c). It disappeared for autobiographical topics where the internet would not help (Experiment 3, p=0.30).
- What it supports
- That access to information is mistaken for knowledge held internally, that the illusion is specific to searchable domains rather than general overconfidence, and that it does not require the search to succeed.
- What it does not support
- That actual explanatory ability changes. Every dependent measure is a self-rating and no experiment tested real knowledge. The authors also note participants were presumably heavier internet users than average.
Peer-reviewed#
Risko, E. F. and Gilbert, S. J. (2016). Cognitive Offloading#
Trends in Cognitive Sciences, 20(9)
- Method
- Review of the experimental literature on offloading.
- Finding
- Defines cognitive offloading as using physical action or an external tool to reduce the mental demand of a task, and shows people offload not only when a task is hard but when they judge it to be hard.
- What it supports
- That the decision to offload is metacognitive, and frequently mistaken.
- What it does not support
- Nothing about generative AI specifically; it predates it.
Peer-reviewed#
Wiradhany, W. and Nieuwenstein, M. R. (2017). Cognitive Control in Media Multitaskers: Two Replication Studies and a Meta-Analysis#
Attention, Perception and Psychophysics, 79(8)
- Method
- Two direct replications of Ophir, Nass and Wagner (2009), 14 tests at mean power 0.81, plus a meta-analysis of 39 effect sizes.
- Finding
- Only five of 14 tests showed increased distractibility, and only two survived a Bayesian analysis. The meta-analytic association became non-significant after correcting for small-study effects. The authors question whether the association exists.
- What it supports
- That one of the most-cited findings about media multitasking and attention does not replicate.
- What it does not support
- That multitasking has no cost at all. This is one specific effect, distractor filtering, not the whole question.
Compiled review#
Phillips, P. J., Hahn, C. A., Fontana, P. C., Broniatowski, D. A. and Przybocki, M. A. (2021). Four Principles of Explainable Artificial Intelligence (NISTIR 8312)#
National Institute of Standards and Technology Interagency Report 8312
- Method
- Framework paper synthesising the explainable-AI literature. Not a measurement study.
- Finding
- Sets out four principles: explanation, meaningful, explanation accuracy and knowledge limits. In the report's own words, explanation accuracy is a distinct concept from decision accuracy, and regardless of the system's decision accuracy the corresponding explanation may or may not accurately describe how the system came to its conclusion. It also notes that the explanation and meaningful principles alone do not require an explanation to reflect the system's actual process.
- What it supports
- That an explanation being intelligible, and even being accurate about the process, is separate from the answer being right. The two are different properties and are routinely treated as one.
- What it does not support
- That explanations help or harm users in practice. This is a framework rather than an experiment, and it reports no effect on human decision quality.
Compiled review#
Porter, T., Elnakouri, A., Meyers, E. A., Shibayama, T., Jayawickreme, E. and Grossmann, I. (2022). Predictors and consequences of intellectual humility#
Nature Reviews Psychology, 1, 524-536
- Method
- Review synthesising definitions and findings on intellectual humility across personality, judgement, education and organisational research.
- Finding
- Identifies a metacognitive core on which there is scholarly consensus, recognising the limits of one's knowledge and being aware of one's fallibility, with social and behavioural features around it: recognising that others may hold legitimate differing beliefs, and willingness to reveal ignorance in order to learn.
- What it supports
- That the metacognitive core is agreed across the field even though the wider construct is not.
- What it does not support
- That the construct is settled, or that it can be measured reliably by asking people. The review notes that where intellectual humility is seen as desirable, self-report makes a false impression easy to create.
Working paper#
Who Goes First? Influences of Human-AI Workflow on Decision Making in Clinical Imaging (2022). Who Goes First? Influences of Human-AI Workflow on Decision Making in Clinical Imaging#
arXiv 2205.09696
- Method
- Between-subjects study, 19 veterinary radiologists reviewing 40 X-rays for 33 findings, comparing seeing the AI output alongside the image against committing to a provisional diagnosis first.
- Finding
- Final diagnoses matched the AI 91 per cent of the time when the AI was seen first, against 89 per cent when the clinician committed first. Where the AI flagged a finding, agreement was 71 per cent against 65 per cent. The anchoring produced only marginal diagnostic gains because of over-reliance on erroneous advice.
- What it supports
- That forming a view before seeing the machine's answer measurably changes the final judgement.
- What it does not support
- Generalisation. Nineteen participants in one clinical speciality, and the effect sizes are small.
Peer-reviewed#
Krügel, S., Ostermaier, A. and Uhl, M. (2023). ChatGPT's inconsistent moral advice influences users' judgment#
Scientific Reports, 13, 4569. DOI 10.1038/s41598-023-31341-0. Published 6 April 2023
- Method
- Preregistered online experiment run on 21 December 2022 with 1,851 US residents recruited through CloudResearch Prime Panels, of whom 767 passed both comprehension checks and form the analysis sample as preregistered. Participants read a transcript of advice on a trolley dilemma, in the switch or the bridge version, arguing for or against sacrificing one life to save five, attributed either to ChatGPT or to a human moral advisor. The advice itself came from ChatGPT, which had given contradictory answers to the same question on 14 December 2022.
- Finding
- The advice moved participants' own moral judgement in both dilemmas, and in the bridge version it flipped the majority verdict. Disclosure made almost no difference: the effect was statistically indistinguishable whether the source was named as a chatbot or as a human advisor. Eighty per cent of participants said they would have reached the same judgement without the advice, and their judgements show they would not have. Only 67 per cent said the same of other participants, and 79 per cent rated themselves more ethical than the others.
- What it supports
- That advice from a model with no settled position still moves the position of the person reading it, that telling them it is a machine does not protect them, and that they cannot see it happening. Transparency, on this evidence, is not a sufficient safeguard.
- What it does not support
- How large the shift is in absolute terms. The paper reports test statistics and figure proportions rather than an effect size in the text, and no confidence intervals appear in the prose. One dilemma type, one sitting, a 41 per cent comprehension pass rate, and a model version from December 2022. It says nothing about repeated real-life decisions or about whether the influence persists.
Peer-reviewed#
Sharma, M., Tong, M., Korbak, T. et al. (2023). Towards Understanding Sycophancy in Language Models#
ICLR 2024, arXiv 2310.13548
- Method
- Analysis of five production AI assistants across four free-form text generation tasks, plus analysis of the human preference datasets used to train them.
- Finding
- All five assistants consistently exhibited sycophancy. Both humans and the preference models trained on their judgements prefer convincingly written sycophantic responses over correct ones a non-negligible share of the time, and optimising against those preference models sometimes sacrifices truthfulness.
- What it supports
- That sycophancy is a predictable consequence of training on human preference, rather than an incidental defect of one product.
- What it does not support
- Any effect on the quality of a user's decisions. It measures model behaviour, not user outcomes.
Peer-reviewed#
Gerlich, M. (2025). AI Tools in Society: Impacts on Cognitive Offloading and the Future of Critical Thinking correction#
Societies, 15(1), 6
- Method
- Survey and interviews, 666 participants.
- Finding
- A negative correlation between frequent AI use and critical-thinking scores, mediated by cognitive offloading, strongest among the youngest users.
- What it supports
- An association, with a plausible mechanism.
- What it does not support
- Causation, and it carries a published correction (Societies 2025, 15(9), 252) which anyone citing it should read alongside.
Compiled review#
Klein, R.M. and Feltmate, B.B.T. (2025). The vigilance decrement: its first 75 years#
Frontiers in Cognition, 4. DOI 10.3389/fcogn.2025.1632885
- Method
- A review of seventy-five years of vigilance research from Mackworth onwards.
- Finding
- The decrement itself has held across the literature. Its mechanism has not: whether the decline reflects falling sensitivity or a shifting response criterion remains disputed, and the sensitivity account has recently been challenged.
- What it supports
- That the effect is durable enough to design around, and that citing it as settled science overstates the position.
- What it does not support
- Any particular mechanism, and therefore any intervention that depends on one. A page arguing that oversight fails for a specific cognitive reason is going beyond what this review supports.
Working paper#
Kosmyna, N. et al. (2025). Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task#
MIT Media Lab preprint, arXiv:2506.08872
- Method
- EEG study, 54 participants, essay writing with an LLM, a search engine, or unaided.
- Finding
- The LLM group showed the weakest brain connectivity and the lowest sense of ownership over their own writing.
- What it supports
- Very little on its own. It is suggestive and widely over-quoted.
- What it does not support
- Anything settled. 54 participants, a preprint, and reproducibility flagged by its own commentators. Treat claims of proof with suspicion.
Peer-reviewed#
Lee, H.-P. et al. (2025). The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects from a Survey of Knowledge Workers#
Microsoft Research and Carnegie Mellon, CHI 2025
- Method
- Survey of 319 knowledge workers about 936 real uses of AI at work.
- Finding
- Higher confidence in the tool was associated with less critical thinking, and the thinking that remains shifts from producing to verifying, from solving to integrating.
- What it supports
- That the character of professional thinking changes with AI use, by self-report.
- What it does not support
- Causation. People who think differently may use AI differently, and a survey cannot separate the two.
Peer-reviewed#
Melumad, S. and Yun, J. H. (2025). Experimental evidence of the effects of large language models versus web search on depth of learning#
PNAS Nexus, 4(10), pgaf316. DOI 10.1093/pnasnexus/pgaf316
- Method
- Seven online and laboratory experiments, four in the paper and three in the supplement. Participants learned a practical topic using either a large language model or web search links, then wrote advice for a friend. Experiment 1: 1,104 participants, real ChatGPT versus real Google. Experiment 2: 1,979 participants, simulated search holding the underlying FACTS identical across conditions. Experiment 3: 250 lab participants, standard Google versus Google with AI Overviews. Advice scored for length, named entities and pairwise similarity.
- Finding
- Experiment 2, with facts held constant: time engaging with results 83.65 seconds with the summary against 124.32 with links; learned new things 3.71 against 3.96; ownership of knowledge 3.41 against 3.66; thought and effort in advice 3.85 against 4.11; advice 64.49 words against 74.22; references to facts 4.00 against 4.61; pairwise cosine similarity between participants' advice 0.224 against 0.072. Rated comprehensiveness did not differ (4.30 against 4.25, p=0.264). A supplementary condition adding real-time web links to the AI summary did not remove the effect, because only 26 per cent of participants clicked any link.
- What it supports
- That the format of a search result, independent of its content, changes how much effort people invest, how deeply they report learning, and the specificity and distinctiveness of what they can then produce.
- What it does not support
- That knowledge objectively declined. Depth of learning is self-reported throughout and no experiment included a recall or comprehension test; time on task is described by the authors as a proxy for effort. Topics were practical how-to tasks over short horizons. The paper states its own total sample twice and inconsistently, as 10,462 in the abstract and 10,426 in the introduction.
Institutional survey#
OpenAI (2025). Sycophancy in GPT-4o: what happened and what we are doing about it#
OpenAI, 29 April and 2 May 2025
- Method
- First-party incident postmortem covering an update released on 25 April 2025 and rolled back from 28 April.
- Finding
- OpenAI attributed the behaviour to weighting short-term user feedback too heavily, which weakened the reward signal that had been holding sycophancy in check. Offline evaluations and A/B tests looked positive. The problem was flagged only by informal qualitative checks, which were overridden.
- What it supports
- That the failure is a measurement failure as much as a training one, and that satisfaction metrics can rise while the product gets worse.
- What it does not support
- Anything independently verified. It is a company's account of its own incident.
Working paper#
Stanković, M., Hirche, E., Kollatzsch, S. and Doetsch, J.N. (2025). Comment on: Your Brain on ChatGPT: Accumulation of Cognitive Debt When Using an AI Assistant for Essay Writing Tasks#
arXiv 2601.00856, 29 December 2025
- Method
- Methodological critique of Kosmyna et al. (arXiv 2506.08872) by researchers at the University of Vienna and TU Dresden. Includes an a priori power analysis using G*Power. Itself a preprint, not peer reviewed.
- Finding
- Argues the MIT cognitive debt study is underpowered: a repeated-measures design at f=0.25, alpha=.05, power=.95 would require approximately 159 participants against the 54 used, with some figures interpreted from subsamples of two to four essays. The strongest objection concerns the construct itself: the search engine group relied on an external tool yet showed no impairment, with the search-engine to brain-only comparison returning p=1, which the authors say contrasts with the interpretation that task delegation increases cognitive debt. Also documents reporting inconsistencies, an unexplained 55 versus 54 participant discrepancy, and unclear FDR correction levels.
- What it supports
- That the most widely cited term in this area rests on a contested pilot, and that the contest is on the record and specific rather than rhetorical.
- What it does not support
- That the MIT findings are wrong. It is a critique offered to improve a manuscript for peer review, it is itself unreviewed, and it does not present competing data.
Peer-reviewed#
Cheng, M., Lee, C., Khadpe, P., Yu, S., Han, D. and Jurafsky, D. (2026). Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence#
Science
- Method
- Eleven models tested against human responses on interpersonal advice, plus two preregistered experiments with 1,604 participants, including a live-interaction study using a real personal conflict.
- Finding
- Models affirmed users' actions about 50 per cent more often than humans did, including 47 per cent endorsement on prompts describing clearly harmful behaviour. Interacting with a sycophantic model reduced participants' willingness to repair an interpersonal conflict and increased their conviction that they were in the right. Participants rated the sycophantic model higher quality, trusted it more, and were more willing to use it again.
- What it supports
- That agreement changes what people subsequently do, and that preference runs in the opposite direction from benefit.
- What it does not support
- An effect on factual or analytical decisions. The scenarios are interpersonal advice, not technical judgement.
Working paper#
Dubois, M., Ududec, C., Summerfield, C. and Luettgau, L. (2026). Ask don't tell: Reducing sycophancy in large language models#
arXiv 2602.23971
- Method
- Factorial experiments on three frontier models using 40 debatable questions rendered in 11 framings, rated over ten epochs, with a follow-up test across 600 personas.
- Finding
- Framing input as a statement rather than a question raised sycophancy by roughly 24 percentage points. Prompting the model to convert a user's statement into a question before answering reduced sycophancy more than instructing it not to be sycophantic.
- What it supports
- That how the user phrases the input matters more than telling the model to behave.
- What it does not support
- That adversarial or devil's advocate prompting works. That was not tested, and no controlled evidence for it was found.
Working paper#
Huemmer, M., Durner, F., Shyiramunda, T. and Cummings-Koether, M.J. (2026). AI, Metacognition, and the Verification Bottleneck: A Three-Wave Longitudinal Study of Human Problem-Solving#
arXiv 2601.17055, 21 January 2026
- Method
- Three-wave longitudinal pilot over six months in an academic setting, convenience sample. Exact sample size is not stated in the abstract. Preprint, no journal reference, no control condition.
- Finding
- Daily AI use rose from 52.4 to 95.7 per cent across the waves. Participants relied most heavily on AI for difficult tasks, 73.9 per cent, while showing declining verification confidence at 68.1 per cent and accuracy of 47.8 per cent on complex tasks. Objective performance fell across problem difficulty from 95.2 to 81.0 to 66.7 to 47.8 per cent, with belief-performance gaps widening to 34.6 percentage points. The authors describe verification rather than solution generation becoming the bottleneck.
- What it supports
- That the verification dimension is measurable and that confidence and accuracy can diverge sharply as task difficulty rises. Useful as a design precedent for measuring verification quality.
- What it does not support
- Any generalisable effect. The authors list their own limitations in the abstract: convenience sampling from a single academic cohort, self-report bias, no control condition, mathematical problems only, and a timeframe too short for skill trajectories. They state causal validation requires randomised trials.
Peer-reviewed#
Yu, S., Cheng, M., Jabbar, A., Sucholutsky, I., Collins, K. M., Jurafsky, D. and Hawkins, R. D. (2026). Cognitive offloading and the speedup illusion in human-AI interaction#
Proceedings of the 48th Annual Meeting of the Cognitive Science Society; also arXiv 2605.23177
- Method
- Preregistered behavioural study, N=1,237, on simple cognitive tasks. Compared forecast completion times against actual completion times, with and without AI assistance, against a control condition in which participants imagined help from another person.
- Finding
- Actual completion times did not differ between independent and AI-assisted completion, while participants predicted AI would be significantly faster. The same bias did not appear when participants imagined help from another person. Participants also reported lower subjective effort with AI at equivalent completion times, so time and effort came apart.
- What it supports
- That people are miscalibrated about AI time savings specifically rather than about assistance in general, and that reported effort is not a proxy for elapsed time.
- What it does not support
- Anything about professional work, output quality or long-horizon tasks. The tasks were short and simple by design, and a forecasting error is not the same thing as a productivity claim.
Learning, practice and expertise
How capability is built, and what removes the building.
Peer-reviewed#
Ericsson, K. A., Krampe, R. T. and Tesch-Romer, C. (1993). The Role of Deliberate Practice in the Acquisition of Expert Performance#
Psychological Review, 100(3), 363-406
- Method
- Two studies of violinists and pianists in Berlin.
- Finding
- Sets out deliberate practice: effortful, targeted activity at the edge of current ability, with feedback, sustained over years.
- What it supports
- That expert performance is built through a specific kind of effortful practice rather than exposure.
- What it does not support
- How much of the difference between performers practice explains. See the Macnamara and Maitra re-examination below.
Peer-reviewed#
McDaniel, M. A., Whetzel, D. L., Schmidt, F. L. and Maurer, S. D. (1994). The Validity of Employment Interviews: A Comprehensive Review and Meta-Analysis#
Journal of Applied Psychology, 79(4)
- Method
- Meta-analysis of 245 validity coefficients from 86,311 individuals.
- Finding
- Structured interviews predict job performance substantially better than unstructured interviews.
- What it supports
- That structure, rather than the interview itself, is what carries the predictive weight.
- What it does not support
- A single stable figure for the gap. Estimates of the size of the structure effect vary considerably between meta-analyses.
Peer-reviewed#
Arthur, W., Bennett, W., Stanush, P. L. and McNelly, T. L. (1998). Factors that influence skill decay and retention: A quantitative review and analysis#
Human Performance, 11(1)
- Method
- Meta-analysis of 189 independent data points from 53 articles.
- Finding
- Skill loss ran from d of -0.01 immediately after training to d of -1.4 after more than 365 days of non-use. Physical, natural and speed-based tasks decayed less than cognitive, artificial and accuracy-based tasks.
- What it supports
- That skill decay is measurable, that it is a function of the interval, and that cognitive skills go first.
- What it does not support
- How fast any particular professional skill decays, or how quickly it can be regained.
Peer-reviewed#
Giedd, J.N., Blumenthal, J., Jeffries, N.O., Castellanos, F.X., Liu, H., Zijdenbos, A., Paus, T., Evans, A.C. and Rapoport, J.L. (1999). Brain development during childhood and adolescence: a longitudinal MRI study#
Nature Neuroscience, 2(10), 861-863
- Method
- Longitudinal MRI, 243 scans from 145 healthy participants, 89 male and 46 female, at the US National Institute of Mental Health.
- Finding
- Cortical grey matter in frontal regions peaks in pre-adolescence and then thins through synaptic pruning, with the prefrontal cortex among the last areas to mature.
- What it supports
- That the prefrontal cortex matures later than other regions. This is real, replicated and not in dispute.
- What it does not support
- Any threshold age. The paper reports a slower trajectory, not an endpoint, and names no age at which development completes. Everything downstream that cites it for a cut-off is citing something it does not contain.
Peer-reviewed#
Maguire, E. A., Gadian, D. G., Johnsrude, I. S., Good, C. D., Ashburner, J., Frackowiak, R. S. J. and Frith, C. D. (2000). Navigation-related structural change in the hippocampi of taxi drivers#
PNAS, 97(8), 4398-4403
- Method
- Cross-sectional structural MRI. 16 right-handed male London taxi drivers with more than 1.5 years driving, against scans of 50 healthy right-handed male non-taxi-drivers.
- Finding
- Posterior hippocampi were significantly larger in taxi drivers than in controls, and hippocampal volume correlated with time spent driving a taxi, positively in the posterior and negatively in the anterior hippocampus.
- What it supports
- That sustained, effortful spatial practice is associated with measurable structural difference in the adult brain, and that the association scales with how long the practice has continued.
- What it does not support
- Causation. It is cross-sectional, so it cannot separate the practice building the brain from a particular brain selecting into the job. It says nothing about what happens when the practice stops, nothing about AI, and nothing about knowledge work.
Peer-reviewed#
Gogtay, N., Giedd, J.N., Lusk, L., Hayashi, K.M., Greenstein, D., Vaituzis, A.C., Nugent, T.F., Herman, D.H., Clasen, L.S., Toga, A.W., Rapoport, J.L. and Thompson, P.M. (2004). Dynamic mapping of human cortical development during childhood through early adulthood#
Proceedings of the National Academy of Sciences, 101(21), 8174-8179
- Method
- A densely sampled subset of THIRTEEN participants from the NIMH longitudinal project, each scanned roughly every two years.
- Finding
- Maps the sequence in which cortical regions mature, with higher-order association cortices maturing after lower-order sensorimotor regions.
- What it supports
- A developmental sequence, in thirteen people.
- What it does not support
- A population age of maturity. Thirteen participants cannot establish one, and the paper does not claim to. This is the study behind the Time Magazine coverage in which the number 25 first appears in public.
Peer-reviewed#
Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T. and Rohrer, D. (2006). Distributed Practice in Verbal Recall Tasks: A Review and Quantitative Synthesis#
Psychological Bulletin, 132(3)
- Method
- Meta-analysis of 839 assessments across 317 experiments in 184 articles.
- Finding
- Spacing and retention interval act jointly. The gap between practice sessions that produces best retention increases as the target retention interval increases.
- What it supports
- That when practice happens changes how much survives, independently of how much practice there is.
- What it does not support
- Application to procedural or professional skill. The synthesis covers verbal recall.
Peer-reviewed#
Roediger, H. L. III and Karpicke, J. D. (2006). Test-Enhanced Learning: Taking Memory Tests Improves Long-Term Retention#
Psychological Science, 17(3), 249-255. DOI 10.1111/j.1467-9280.2006.01693.x
- Method
- Two experiments, 120 and 180 participants, reading prose passages on general science topics. Restudying compared with being tested, with final recall measured at 5 minutes, 2 days or 1 week.
- Finding
- The winner reverses with delay. At 5 minutes restudying beat testing, 81 per cent against 75 per cent. At one week testing beat restudying, 56 per cent against 42 per cent. In the second experiment repeated study led at 5 minutes, 83 against 71 per cent, and trailed badly at one week, 40 against 61 per cent.
- What it supports
- That the study method producing the best immediate performance produces the worst durable retention, and that a measurement taken close to the learning will rank the methods in exactly the wrong order.
- What it does not support
- Anything about AI. It is prose recall in a laboratory, and the transfer to professional judgement is by analogy.
Peer-reviewed#
Kapur, M. (2008). Productive Failure#
Cognition and Instruction, 26(3), 379-424. DOI 10.1080/07370000802212669
- Method
- Randomised comparison, 309 eleventh-grade physics students in India working on Newtonian kinematics. One group solved ill-structured problems in groups before individual well-structured problems; the other solved well-structured problems throughout.
- Finding
- The group given ill-structured problems struggled visibly and produced poor solutions during the collaborative phase, then outperformed the other group on individual near-transfer and far-transfer measures afterwards.
- What it supports
- That struggle which looks like failure at the time can produce better subsequent transfer than a smooth path through the same material, and that judging a learning design by how well it is going is unreliable.
- What it does not support
- A precise magnitude. The design, sample and direction of the result are confirmed from the publisher's abstract and the author's own presentation of the study, but the full text sits behind a paywall that could not be read, so no post-test figures are quoted here.
Compiled review#
American Red Cross Advisory Council on First Aid, Aquatics, Safety and Preparedness (2009). Scientific Review: CPR Skill Retention#
American Red Cross
- Method
- Systematic review of 47 articles on CPR skill retention across healthcare and lay populations, with retest intervals from six weeks to 24 months.
- Finding
- Substantial skill degradation occurs within the first year after training, with declining retention from six to twelve months unless there is refresher training.
- What it supports
- That a life-critical, heavily trained procedural skill decays on a timescale of months without practice.
- What it does not support
- Any link to patient outcomes. Decay was measured on manikins, not in resuscitations.
Peer-reviewed#
Bjork, E. L. and Bjork, R. A. (2011). Making Things Hard on Yourself, But in a Good Way: Creating Desirable Difficulties to Enhance Learning#
In Psychology and the Real World, Worth Publishers
- Method
- Synthesis of decades of laboratory work on spacing, interleaving and retrieval practice.
- Finding
- Conditions that make study feel harder improve long-term retention; conditions that make it feel fluent improve immediate performance and worsen retention. Learners systematically mistake fluency for learning.
- What it supports
- That the subjective sense of learning is an unreliable guide to whether learning occurred. Directly relevant, because AI makes work feel fluent.
- What it does not support
- That AI-assisted work is equivalent to a fluent study condition. That inference is ours, not the authors'.
Peer-reviewed#
Karpicke, J. D. and Blunt, J. R. (2011). Retrieval Practice Produces More Learning than Elaborative Studying with Concept Mapping#
Science, 331(6018), 772-775. DOI 10.1126/science.1199327
- Method
- Two experiments, 80 and 120 undergraduates, comparing retrieval practice with elaborative study by concept mapping, tested one week later.
- Finding
- Retrieval practice scored 0.67 against 0.45 for concept mapping, about a 50 per cent advantage in long-term retention, d = 1.50. In the second experiment 101 of 120 students, 84 per cent, did better after retrieval practice than after elaborative study.
- What it supports
- That effortful recall outperforms a more elaborate and more comfortable-feeling study method, on the same material, in the same students.
- What it does not support
- That concept mapping is worthless. A published comment by Mintzes and colleagues (Science, 2011, 334(6055), 453) disputes the instructional fidelity of the concept-mapping condition, and anyone citing this should note it.
Peer-reviewed#
Woollett, K. and Maguire, E. A. (2011). Acquiring 'the Knowledge' of London's Layout Drives Structural Brain Changes#
Current Biology, 21(24), 2109-2114, 20 December 2011
- Method
- Longitudinal structural MRI over four years. 79 male trainee London taxi drivers studying for the Knowledge, and 31 male non-taxi-driver controls, scanned before and after. Average-IQ adults, real training rather than a laboratory task.
- Finding
- In those who qualified, acquiring an internal spatial representation of London was associated with a selective increase in grey matter volume in the posterior hippocampi, with concomitant changes to their memory profile. In the authors' words, no structural brain changes were observed in trainees who failed to qualify or in control participants. The gain in posterior hippocampus came alongside costs elsewhere in the memory profile.
- What it supports
- The strongest available evidence that the practice itself does the work rather than selection. Same starting cohort, same training, and the structural change appears only in those who completed it. It also shows the trade: the capability gained is paid for with capability elsewhere, which is skill substitution observed in tissue rather than argued.
- What it does not support
- Anything about AI, and anything about removal. This is acquisition, over four years, in one domain. It does not show that the structure regresses when the practice is delegated to a machine, and no study has shown that. The frequently circulated claim that GPS use produces measurable cognitive decline or early-onset dementia is a prediction rather than a finding, and cannot be traced to a study of this kind.
Institutional survey#
Bureau d'Enquetes et d'Analyses (2012). Final Report on the accident on 1 June 2009 to the Airbus A330-203, flight AF 447#
BEA, France
- Method
- Statutory accident investigation. Flight recorder analysis and training record review. 25 new safety recommendations.
- Finding
- The investigation cited the lack of practical training in high-altitude manual handling and in the procedure for speed anomalies among the contributing factors.
- What it supports
- That a formal investigation attributed part of an accident to manual handling practice that had not been maintained.
- What it does not support
- A general rate of skill loss across the pilot population. It is one accident.
Peer-reviewed#
Moonen-van Loon, J. M. W., Overeem, K., Donkers, H. H. L. M., van der Vleuten, C. P. M. and Driessen, E. W. (2013). Composite reliability of a workplace-based assessment toolbox for postgraduate medical education#
Advances in Health Sciences Education, 18(5)
- Method
- Generalisability study of 12,779 workplace-based assessments from 953 medical residents.
- Finding
- A reliability coefficient of 0.80 required eight mini-CEX observations, nine DOPS or nine multi-source feedback rounds. Combined in a portfolio the requirement fell to seven, eight and one respectively.
- What it supports
- That observing someone at work can reach defensible reliability, and roughly how many observations that takes.
- What it does not support
- That these scores predict later performance or patient outcomes. It measures consistency, not criterion validity. A single observation is not reliable.
Institutional modelling#
PARC/CAST Flight Deck Automation Working Group (2013). Operational Use of Flight Path Management Systems: Final Report#
Federal Aviation Administration
- Method
- Working-group synthesis of accident and incident data, operator surveys and prior research. 28 findings, 18 recommendations.
- Finding
- Identified vulnerabilities in manual handling after transition from automated control, and in the definition, development and retention of those skills. Also found that pilots sometimes rely too much on automated systems and may be reluctant to intervene.
- What it supports
- That a regulator examined automation dependence in a whole industry and named skill retention as a finding.
- What it does not support
- A measured rate of skill decay. It is a synthesis of findings, not a controlled study.
Peer-reviewed#
Casner, S. M., Geven, R. W., Recker, M. P. and Schooler, J. W. (2014). The Retention of Manual Flying Skills in the Automated Cockpit#
Human Factors, 56(8)
- Method
- 16 airline pilots flew routine and non-routine scenarios in a Boeing 747-400 simulator with automation level varied.
- Finding
- Instrument scanning and manual control were mostly intact even where pilots reported little recent practice. The cognitive tasks, tracking position without a map, deciding the next navigational step and recognising instrument failures, showed frequent and significant problems.
- What it supports
- That the hands survive disuse better than the judgement does, in the profession with the most automation experience.
- What it does not support
- A general rule for knowledge work. The sample is 16 pilots in a simulator.
Compiled review#
Macnamara, B. N., Hambrick, D. Z. and Oswald, F. L. (2014). Deliberate Practice and Performance in Music, Games, Sports, Education, and Professions: A Meta-Analysis#
Psychological Science, 25(8), 1608-1618
- Method
- Meta-analysis of 88 studies and 157 effect sizes relating accumulated deliberate practice to performance.
- Finding
- Deliberate practice explained 12 per cent of the variance in performance overall, 95% CI [9%, 15%], leaving 88 per cent unexplained. By domain: 26 per cent for games, 21 per cent for music, 18 per cent for sports, 4 per cent for education and under 1 per cent for the professions.
- What it supports
- That accumulated practice is one contributor among several rather than the dominant explanation of expert performance, and that its contribution varies sharply by domain.
- What it does not support
- That practice does not build professional expertise. The professions estimate rests on 7 effect sizes, is not statistically significant (p = .62), and the occupations sampled were computer programming, military aircraft piloting, soccer refereeing and insurance selling. It is not evidence about law, medicine, consulting or analysis. It gets quoted as though it were.
Peer-reviewed#
Rowland, C. A. (2014). The Effect of Testing Versus Restudy on Retention: A Meta-Analytic Review of the Testing Effect#
Psychological Bulletin, 140(6), 1432-1463. DOI 10.1037/a0037559
- Method
- Meta-analysis, 159 effect sizes from 61 studies published between 1975 and 2013, random-effects model.
- Finding
- A reliable testing effect, g = 0.50 with a confidence interval of 0.42 to 0.58. With feedback the effect rises to g = 0.73; without feedback it falls to g = 0.39.
- What it supports
- That the testing effect survives aggregation across four decades of studies, at a moderate effect size, and that feedback roughly doubles it.
- What it does not support
- That the effect size transfers to workplace learning. The constituent studies are overwhelmingly laboratory and classroom work on verbal material.
Peer-reviewed#
Hartshorne, J.K. and Germine, L.T. (2015). When does cognitive functioning peak? The asynchronous rise and fall of different cognitive abilities across the life span#
Psychological Science, 26(4), 433-443
- Method
- Cross-sectional analysis of large web-based and standardisation samples across a wide age range.
- Finding
- Different cognitive abilities peak at different ages, some in the late teens or early twenties, others not until the forties or fifties. There is no single age at which cognitive functioning peaks.
- What it supports
- That a single maturity age is the wrong shape of answer, whatever number is put in it. Abilities do not arrive together.
- What it does not support
- That age is irrelevant. It shows the timing is ability-specific rather than absent.
Peer-reviewed#
Murre, J. M. J. and Dros, J. (2015). Replication and Analysis of Ebbinghaus' Forgetting Curve#
PLoS ONE, 10(7)
- Method
- Single-subject relearning experiment replicating Ebbinghaus across intervals from 20 minutes to 31 days, 10 replications per interval.
- Finding
- Relearning to criterion took less time than original learning at every retention interval tested, confirming the savings effect Ebbinghaus reported in 1885.
- What it supports
- That something survives apparent forgetting, and that it shows up as faster relearning rather than as recall.
- What it does not support
- That professional skills behave like nonsense syllables, or that savings hold at the scale of a career. It is one subject and verbal material.
Peer-reviewed#
Storm, B. C. and Stone, S. M. (2015). Saving-Enhanced Memory: The Benefits of Saving on the Learning and Remembering of New Information#
Psychological Science, 26(2), 182-188. DOI 10.1177/0956797614559285
- Method
- Three experiments on the effect of saving a digital file on memory for subsequently studied material. GRADED FROM THE PUBLISHED ABSTRACT ONLY: the full text is paywalled and could not be read at the primary source, so no sample sizes or statistics are recorded here.
- Finding
- Saving one file before studying a new file significantly improved memory for the contents of the new file. The effect was not observed when the saving process was deemed unreliable, or when the contents of the to-be-saved file were not substantial enough to interfere with memory for the new file.
- What it supports
- That cognitive offloading can improve rather than degrade subsequent memory, and that the benefit depends on the external store being trusted.
- What it does not support
- Anything quantified. This entry rests on the abstract alone, which is a weaker basis than every other entry in this base and is stated as such. It also predates generative AI and concerns file saving rather than a system that can fabricate its own contents.
Compiled review#
Simons, D. J., Boot, W. R., Charness, N., Gathercole, S. E., Chabris, C. F., Hambrick, D. Z. and Stine-Morrow, E. A. L. (2016). Do Brain-Training Programs Work?#
Psychological Science in the Public Interest, 17(3)
- Method
- Systematic review applying pre-specified best-practice standards to every study cited by commercial brain-training companies as evidence of efficacy.
- Finding
- Extensive evidence that training improves performance on the trained tasks, less evidence for closely related tasks, and little evidence that training improves distantly related tasks or everyday cognitive performance. No cited study met all best-practice standards.
- What it supports
- That near transfer is real and far transfer, which is what any claim to have trained attention requires, is essentially unsupported.
- What it does not support
- That practising a specific skill is useless. The failure is transfer, not learning.
Peer-reviewed#
Przybylski, A. K. and Weinstein, N. (2017). A Large-Scale Test of the Goldilocks Hypothesis: Quantifying the Relations Between Digital-Screen Use and the Mental Well-Being of Adolescents#
Psychological Science, 28(2), 204-215
- Method
- Preregistered analysis of a representative sample of English 15-year-olds. Sampling frame 298,080; 120,115 provided usable data, 100,850 on paper and 19,265 online.
- Finding
- Links between digital screen time and mental wellbeing are described by quadratic rather than linear functions, with inflection points at 1 hour 40 minutes for weekday video-game play and 1 hour 57 minutes for weekday smartphone use, rising to 3 hours 41 minutes and 4 hours 17 minutes for other measures and to between 3 hours 35 minutes and 4 hours 50 minutes at weekends. Average Cohen's d for engagement beyond the inflection points was minus 0.18, accounting for 1 per cent or less of variance, against d = 0.54 for regularly eating breakfast and 0.58 for regular sleep. The authors state that moderate use is not intrinsically harmful and may be advantageous.
- What it supports
- That dose-response in this literature is not linear, and that the negative effects of heavy use are less than a third the size of the positive associations with sleep and breakfast.
- What it does not support
- Causation, and not displacement. The paper names the displacement hypothesis as the field's dominant assumption and calls for future work systematically analysing what is being displaced or amplified, which the field did not then do. Cross-sectional, self-reported exposure, English 15-year-olds only.
Compiled review#
National Academies of Sciences, Engineering, and Medicine (2018). How People Learn II: Learners, Contexts, and Cultures#
The National Academies Press, Washington DC
- Method
- Consensus study report synthesising research on learning across the lifespan.
- Finding
- Defines metacognition as the ability to monitor and regulate one's own cognitive processes and to consciously regulate behaviour, including affective behaviour, and identifies calibration as the accuracy of a learner's monitoring.
- What it supports
- That the established definition contains a control component as well as a knowledge component. Monitoring alone is not metacognition.
- What it does not support
- Anything about AI. The report predates general availability of these systems and makes no claim about them.
Peer-reviewed#
Macnamara, B. N. and Maitra, M. (2019). The role of deliberate practice in expert performance: revisiting Ericsson, Krampe and Tesch-Romer (1993)#
Royal Society Open Science, 6, 190327
- Method
- Direct replication and re-analysis of the 1993 study.
- Finding
- Accumulated practice explained considerably less of the difference between performers than the original is usually taken to claim.
- What it supports
- That the quality and design of practice matters more than the count. Included here deliberately, because it complicates the argument this research relies on.
- What it does not support
- That practice does not matter. It does; the simple dose-response reading is what fails.
Peer-reviewed#
Orben, A. and Przybylski, A. K. (2019). The association between adolescent well-being and digital technology use#
Nature Human Behaviour, 3(2), 173-182
- Method
- Specification curve analysis across three nationally representative datasets: the US Youth Risk and Behaviour Survey 2007-2015 at 74,814 adolescents, Monitoring the Future 2008-2016 at 268,672, and the UK Millennium Cohort Study at 11,872, a total of 355,358. 372 justifiable specifications identified for YRBS, 40,966 for MTF and 603,979,752 for MCS, of which 20,004 were run.
- Finding
- The association between digital technology use and adolescent wellbeing is negative but small, explaining at most 0.4 per cent of the variation in wellbeing, which the authors state is too small to warrant policy change. In YRBS, regularly eating potatoes was associated with wellbeing 0.9 times as negatively as technology use; in MCS, wearing glasses was 1.5 times as negatively. Bullying ran 4.3 times more negative and marijuana 2.7 times, both YRBS. Sleep and breakfast ranged from 1.7 to 44.2 times more positive across all three datasets.
- What it supports
- That an entire public debate was conducted on an effect too small to act on, and that analytic flexibility rather than data availability was doing the work in the studies that found more.
- What it does not support
- Causation in either direction. The authors state it is possible the associations they document, and those previously documented, are spurious, and that noisy self-report measurement could itself have diminished a real effect. Two figures often attached to this paper are not in it: 1.45 times for glasses, and 3,221,225,472 analyses, which comes from Odgers and Jensen describing it. Read in the Oxford ORA accepted manuscript, the published full text being paywalled.
Peer-reviewed#
Orben, A. and Przybylski, A. K. (2019). Screens, Teens, and Psychological Well-Being: Evidence From Three Time-Use-Diary Studies#
Psychological Science, 30(5), 682-696
- Method
- Exploratory and confirmatory specification analyses across three nationally representative datasets from Ireland, the United States and the United Kingdom, N = 17,247 after exclusions, using time-use diaries as well as retrospective self-report.
- Finding
- Little evidence of substantial negative associations between digital screen engagement and adolescent wellbeing, whether measured across the day or before bedtime. Correlations between diary-recorded and retrospectively self-reported engagement were 0.18 in Ireland, 0.08 and 0.05 on US weekdays and weekend days, and 0.18 in the UK. Retrospective self-report consistently produced the most negative correlations. Extrapolating from median effects, an adolescent would need to report 63 hours 31 minutes more technology use a day to lower wellbeing by half a standard deviation, or 11 hours 14 minutes taking the maximum effect size in the specification set.
- What it supports
- That the two standard instruments for measuring screen exposure barely agree with each other, and that the more negative results come from the weaker instrument, which is what common method variance would predict.
- What it does not support
- That diaries are ground truth. They remain recall-based and the authors note brief or concurrent uses may not be recorded. Still cross-sectional, and it measures totals and timing rather than the kind of activity displaced.
Peer-reviewed#
Dahmani, L. and Bohbot, V. D. (2020). Habitual use of GPS negatively impacts spatial memory during self-guided navigation#
Scientific Reports, 10, 6310, 14 April 2020. DOI 10.1038/s41598-020-62877-0. Read in full at source 5 September 2026
- Method
- Behavioural, cross-sectional with an unplanned longitudinal arm. 50 healthy regular drivers in Montreal aged 19 to 35 (18 women, 32 men; mean age 27.6), driving at least four days a week, from 60 recruited. Lifetime GPS experience measured by the McGill GPS questionnaire; navigation measured on two virtual radial-arm mazes plus a map-drawing score and the Santa Barbara Sense of Direction scale. 13 of the 50 returned a mean of 3.23 years later. Effect sizes are Pearson r with bootstrapped one-tailed BCa 95 per cent intervals. NO NEUROIMAGING: hippocampal dependence is inferred from prior validation of the tasks, not measured in this sample.
- Finding
- Cross-sectionally, greater lifetime GPS experience was associated with lower use of hippocampus-dependent spatial strategies (r = -0.22 on the first probe trial), lower navigation strategy scores (r = -0.20), poorer map drawing (r = -0.22) and fewer landmarks noticed (r = -0.26). In the 13-person follow-up, hours of GPS use since first testing tracked a steeper decline in spatial memory strategy use (r = -0.68) and in map drawing (r = -0.52).
- What it supports
- The only study to follow the same people while their GPS use rose. Its strongest internal argument against reverse causation is that heavier GPS users did not report a poorer sense of direction (r = 0.07 against the SBSOD), so the obvious alternative, that weak navigators reach for the satnav, has no support in the data.
- What it does not support
- Anything about the brain, because no scan was taken. Anything about dementia, Alzheimer's or atrophy: those words appear nowhere in the paper. And little with confidence about the longitudinal effect, which rests on 13 people from an unplanned follow-up. The authors' own words: they caution against any strong conclusions as spurious correlations are possible. Their discussion elsewhere uses notably firmer causal language than that sentence licenses. Nothing here transfers to reasoning: spatial memory is not judgement.
Compiled review#
Odgers, C. L. and Jensen, M. R. (2020). Annual Research Review: Adolescent mental health in the digital age: facts, fears, and future directions#
Journal of Child Psychology and Psychiatry, 61(3), 336-348
- Method
- Annual research review of the evidence on adolescent digital technology use and mental health, covering 29 studies in its main table.
- Finding
- Most research to date has been correlational, focused on adults rather than adolescents, and has produced a mix of conflicting small positive, negative and null associations. The most recent and rigorous large-scale preregistered studies report small associations that offer no way of distinguishing cause from effect and are unlikely to be of clinical or practical significance, explaining less than 0.5 per cent of the variance. Of the 29 studies reviewed, only two included objective or informant-rated measures of screen use, and the correlation between objectively measured and retrospectively reported screen time is estimated at about 0.20.
- What it supports
- That the field's own review verdict is that the evidence does not support causal claims or even strong consistent correlational patterns.
- What it does not support
- Anything new. It is a review, not primary data, it does not test displacement of any specific activity, and it does not address AI. Read in the NIH author manuscript.
Peer-reviewed#
Parry, D. A., Davidson, B. I., Sewall, C. J. R., Fisher, J. T., Mieczkowski, H. and Quintana, D. S. (2021). A systematic review and meta-analysis of discrepancies between logged and self-reported digital media use#
Nature Human Behaviour, 5(11), 1535-1547
- Method
- Systematic review and meta-analysis using robust variance estimation. 106 effect sizes included overall; the self-report against logged comparison draws on 66 effect sizes from 44 studies with a total sample of 52,007.
- Finding
- The correlation between self-reported and logged digital media use is positive but medium, r = 0.38, 95 per cent CI 0.33 to 0.42. For problematic use it falls to r = 0.25 across 40 effect sizes from 19 studies. Over-reporting and under-reporting occur in similar proportions, and fewer than 10 per cent of self-reports fall within 5 per cent of the equivalent logged value. The authors conclude that self-report measures may not be a valid stand-in for more objective measures and ask for pause in drawing wide-reaching knowledge or policy conclusions from studies relying solely on them.
- What it supports
- That the independent variable in most of the screen-time literature is wrong by a measured amount, which is the strongest single reason not to carry that method into questions about AI use.
- What it does not support
- That self-report is useless, or that logs are ground truth: the authors note potential biases in log data too. They state it is an open question whether the discrepancy is random or systematic error. Nothing in it concerns AI. Read in the University of Bath accepted manuscript alongside the published abstract.
Peer-reviewed#
Sinha, T. and Kapur, M. (2021). When Problem Solving Followed by Instruction Works: Evidence for Productive Failure#
Review of Educational Research, 91(5), 761-798. DOI 10.3102/00346543211019105
- Method
- Meta-analysis of 53 studies and 166 comparisons of problem-solving-before-instruction against instruction-before-problem-solving.
- Finding
- A moderate effect favouring problem solving first, Hedges g = 0.36 with a confidence interval of 0.20 to 0.51, rising to between 0.37 and 0.58 where the design followed the productive failure principles closely. The effect reverses for second to fifth graders and for domain-general skills, where instruction first wins.
- What it supports
- That letting people struggle before teaching them beats teaching them first, at moderate effect size, for older learners on domain-specific content.
- What it does not support
- That struggle is universally good. The authors report the reversal for young children and for general skills themselves, and an earlier meta-analysis by Darabi and colleagues rested on only 12 studies.
Compiled review#
Verhaeghen, P. (2021). Mindfulness as Attention Training: Meta-Analyses on the Links Between Attention Performance and Mindfulness Interventions, Long-Term Meditation Practice, and Trait Mindfulness#
Mindfulness, 12(3)
- Method
- Three meta-analyses covering 109 effect sizes from 40 intervention studies, 59 effect sizes from 18 long-term meditator studies, and 197 effect sizes from 28 trait studies.
- Finding
- Average effects were small to moderate, Hedges g of 0.29 for interventions and 0.32 for long-term practice, concentrated in inhibition and executive control rather than sustained attention.
- What it supports
- That something measurable happens, and that it is smaller and narrower than the popular claim.
- What it does not support
- That the effect survives comparison with an active control. The published breakdown does not separate active from passive controls, so demand effects cannot be excluded.
Peer-reviewed#
Brynjolfsson, E., Li, D. and Raymond, L. (2023). Generative AI at Work#
Quarterly Journal of Economics, 140(2), 889-942. DOI 10.1093/qje/qjae044. Advance Access 4 February 2025. Earlier version NBER Working Paper 31161
- Method
- Staggered rollout of a GPT-3-based conversational assistant across 5,172 customer-support agents in 133 teams at a single Fortune 500 business-process software firm, most of them working from the Philippines. Three million chats observed, 1.2 million of them post-deployment. The system was fine-tuned on past agent conversations, with chats by top performers deliberately up-weighted in training.
- Finding
- Resolutions per hour rose 15 per cent on average, 15.2 per cent in the preferred specification with agent and tenure fixed effects. Less skilled and less experienced workers gained a 30 per cent increase in issues resolved per hour, rising to 36 per cent for the lowest skill quintile, while the most skilled saw no significant productivity change and small declines in conversation quality and customer satisfaction. Customer sentiment improved by half a standard deviation and requests to speak to a manager fell about 25 per cent. During unplanned outages, agents with longer AI exposure still handled chats faster than their pre-AI baseline, but only those who had adhered closely to the suggestions.
- What it supports
- That a model trained on the behaviour of a firm's best workers can transfer measurable parts of that behaviour to its newest ones, immediately and at scale, raising the floor far more than the ceiling. The outage evidence shows some of the gain survives the tool being switched off, conditional on the worker having engaged with it rather than passed it through.
- What it does not support
- Whether those novices became experts. It measures output over months in one firm, one occupation and a stable product environment, and the authors say so. Wages, labour demand and hiring composition were not observed. The outage estimates are the authors' own noisiest, because outages are rare and may not be comparable chats.
Peer-reviewed#
Liang, W., Yuksekgonul, M., Mao, Y., Wu, E. and Zou, J. (2023). GPT detectors are biased against non-native English writers#
Patterns, 4(7)
- Method
- Seven widely used GPT detectors evaluated against TOEFL essays by non-native English speakers and essays by US eighth-grade students.
- Finding
- Detectors misclassified more than half of the non-native essays as AI-generated, an average false positive rate of 61.22 percent, while classifying US eighth-grade essays with near-perfect accuracy. The proposed mechanism is that detectors rely on perplexity, and second-language writing is more predictable.
- What it supports
- That AI detection carries a severe and systematic bias against second-language writers, and that the bias is structural rather than a tuning problem.
- What it does not support
- That every detector now on the market performs identically. The study tested tools available at the time, and vendors dispute the generalisation.
Peer-reviewed#
Sackett, P. R., Zhang, C., Berry, C. M. and Lievens, F. (2023). Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors#
Industrial and Organizational Psychology, 16
- Method
- Meta-analytic re-correction of prior personnel-selection meta-analyses, addressing systematic overcorrection for range restriction.
- Finding
- Work sample validity falls from the widely quoted .54 to .33. Structured interviews fall from .51 to .42 and become the strongest single predictor. Unstructured interviews fall from .38 to .19. General cognitive ability falls from .51 to .31.
- What it supports
- That the numbers most often cited for assessment methods were inflated, and by how much.
- What it does not support
- That these methods do not work. Work samples and structured interviews remain among the strongest predictors available. It also offers no validity data for AI-era assessment formats.
Peer-reviewed#
Adinoff, B. and Nunes, J.C. (2025). Challenging the 25-year-old 'mature brain' mythology: implications for the minimum legal age for non-medical cannabis use#
The American Journal of Drug and Alcohol Abuse, 51(5), 577-583
- Method
- Review of the neuroscience and policy literature behind the age-25 threshold.
- Finding
- Argues the mature-brain-at-25 claim is not supported by the underlying neuroscience and should not be used as a basis for age thresholds in policy.
- What it supports
- That the challenge to this claim is in the peer-reviewed literature rather than confined to science journalism.
- What it does not support
- Anything about AI, learning or capability. It is cited here for the status of the claim, not for the subject.
Peer-reviewed#
Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakci, O. and Mariman, R. (2025). Generative AI Without Guardrails Can Harm Learning: Evidence from High School Mathematics#
Proceedings of the National Academy of Sciences, 122(26)
- Method
- Field experiment, nearly 1,000 high-school students, three arms: unrestricted GPT-4, a hints-only tutor, and a control.
- Finding
- Grades rose 48 percent with unrestricted access and 127 percent with the tutor while the tool was present. With access removed, the unrestricted group scored 17 percent LOWER than students who never had it. The guardrailed tutor largely removed the harm.
- What it supports
- That the design of the interface, not the presence of AI, decides whether people learn. The single most useful result in this literature.
- What it does not support
- What a guardrailed interface should look like for professional work. It was school mathematics over a bounded period.
Peer-reviewed#
Kestin, G., Miller, K., Klales, A., Milbourne, T. and Ponti, G. (2025). AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting#
Scientific Reports, 15, published 3 June 2025
- Method
- Randomised crossover experiment in Harvard's largest introductory physics course, Fall 2023. Of 233 enrolled students, 194 were eligible on consent and completion. Each student experienced both conditions across two consecutive weeks, one topic taught by in-class active learning and one by a purpose-built AI tutor at home, with pre-tests and post-tests for each. The tutor used GPT-4 with expert-crafted question-specific prompts, pre-written answers, instructional video and a structured scaffold.
- Finding
- Median post-test score 4.5 in the AI condition against 3.5 in the active-learning condition, from a combined pre-test median of 2.75; median learning gain over double. Mann-Whitney z = -5.6, p below 10 to the minus 8. Linear regression effect size 0.63, described by the authors as an underestimate because of a ceiling effect; quantile regression gives 0.73 to 1.3 standard deviations. Median time on task 49 minutes against 60 assumed for the class, with no correlation between time on task and score. Engagement 4.1 against 3.6 and motivation 3.4 against 3.1; enjoyment and growth mindset showed no significant difference.
- What it supports
- That a heavily engineered AI tutor, built by subject experts to follow established pedagogy, can outperform a well-run active-learning class on immediate post-test performance at the understanding, applying and analysing levels, in less time.
- What it does not support
- That a general chatbot does this. Accuracy depended on pre-written answers and instructor-written prompts. Retention was not measured; post-tests followed the lessons immediately. One course, one institution, two topics. The authors state they do not presume the result holds where complex synthesis or higher-order critical thinking is required.
Compiled review#
Mansfield, K. L., Ghai, S., Hakman, T., Ballou, N., Vuorre, M. and Przybylski, A. K. (2025). From social media to artificial intelligence: improving research on digital harms in youth#
The Lancet Child and Adolescent Health, 9(3), 194-204. Personal View, not primary research
- Method
- Personal View setting out the methodological failures of the social media harms literature and what should be done differently for AI. No new data.
- Finding
- Self-reported screen time is problematic as a measure, being imprecise and prone to bias, and as a construct, being unidimensional, homogenous and of little validity, failing to distinguish social, educational, entertainment, work and informational uses. On AI, the authors write that using self-reported frequency or duration of adolescent AI use as the exposure measure of interest is perhaps even more concerning than counting total time spent on social media, and that only behavioural data on exposure to a range of AI applications would provide the detail needed. They record that health policy decisions have been implemented on inconsistent, non-causal or ungeneralisable evidence of online harms.
- What it supports
- That the researchers who built the screen-time evidence base have published the warning against transplanting its method to AI, in advance and in a clinical journal.
- What it does not support
- Anything empirical. It is a commentary with no new data and no head-to-head methodological comparison. Its prescription is finer-grained measurement of which AI in what context, which is adjacent to but not the same as asking which human practice was displaced. Publisher full texts returned empty bodies; read in the Oxford ORA accepted version.
Peer-reviewed#
Mousley, A. and colleagues (2025). Topological turning points across the human lifespan#
Nature Communications, November 2025
- Method
- Diffusion MRI from 3,802 people aged 0 to 90, analysed for structural network topology across the lifespan.
- Finding
- Four topological turning points, at approximately ages 9, 32, 66 and 83, dividing life into five epochs. The adolescent epoch runs from about 9 to about 32. The largest overall shift in trajectory occurs around 32, not in the mid-twenties.
- What it supports
- That structural network reorganisation continues well past the mid-twenties, and that the nearest thing to a boundary at the end of adolescence sits around 32.
- What it does not support
- That 32 is the new 25. The authors describe turning points in network topology, not a moment of cognitive completion, and reading it as a new threshold would repeat the original error with a different number.
Peer-reviewed#
Prasad, S. et al. (2025). Enhancing medical assessment strategies: a comparative study between structured, traditional and hybrid viva-voce assessment#
BMC Medical Education, 25
- Method
- 151 medical students assessed by two examiners across three viva formats, compared on reliability and perceived fairness.
- Finding
- Traditional unstructured viva showed significant inter-examiner variability. Structured formats improved fairness and coverage. The best format reached a reliability of 0.663, which is moderate rather than high.
- What it supports
- That oral examination can be made fairer by structuring it, and that unstructured viva has a measurable examiner problem.
- What it does not support
- That oral assessment is highly reliable. Even the best format tested was only moderate, at a single institution.
Working paper#
Cruces, G., Fernandez Meijide, D., Galiani, S., Galvez, R. and Lombardi, M. (2026). Does generative AI narrow education-based productivity gaps? Evidence from a randomized experiment#
NBER Working Paper 34851; CEPR Discussion Paper 21299; arXiv:2608.04198
- Method
- Randomised online experiment, 1,174 adults aged 25 to 45, workplace-style problem-solving task with or without a generative AI assistant, followed by an unassisted module.
- Finding
- AI improved performance for everyone and more for the less educated. Without AI, higher-education participants outperformed lower-education participants by 0.548 standard deviations; with AI the gap fell to 0.139, closing about three-quarters of it. Treated participants did not perform worse once AI was removed, and lower-education participants retained part of their improvement, although a sizeable gap re-emerged.
- What it supports
- That assisted use does not automatically leave people worse off than unassisted controls when the tool is taken away, and that AI can compress an education-based performance gap while it is present.
- What it does not support
- That skill formed. One session with an immediate unassisted module tests transfer within a sitting, not skill formation over time, and the studies that found post-removal deficits taught a body of knowledge and removed the tool afterwards. The equity gain is also partly transient by the authors' own account, since a sizeable gap re-emerges without the assistant.
Argued perspective#
Ke, Y., Jin, L., Ong, J.C.L., Thirunavukarasu, A.J., Car, J., Cheung, C.Y., Tham, Y.C., Ting, D.S.W., Ong, M.E.H., Compton, S., Narayan, A., Keane, P.A., Wong, T.Y., Bates, D.W., Tan, P. and Liu, N. (2026). AI-induced never-skilling in medical education#
Nature Medicine, 32(6), 1997-2006, published 22 May 2026. DOI 10.1038/s41591-026-04438-y
- Method
- A Perspective, not primary research. Sixteen authors across sixteen institutions argue a three-part taxonomy of how AI can interrupt the formation of clinical competence, and propose an untested three-phase protective framework. The clinical evidence it leans on is borrowed, chiefly Budzyn and colleagues on colonoscopy.
- Finding
- Separates three distinct failures. Deskilling is the degradation of established competence in clinicians already trained. Mis-skilling is the acquisition of incorrect reasoning patterns through uncritical adoption of erroneous or biased AI output. Never-skilling is the failure to form foundational competence during training, when AI substitutes for the cognitive effort that would have built it. The authors predict a state they call false proficiency: competence that appears real but depends on the AI remaining available.
- What it supports
- That the three failures are different problems requiring different responses, and that the entry-level case is not simply deskilling applied to younger people. Never-skilling has no baseline to return to, which is what makes it a distinct category rather than a matter of degree.
- What it does not support
- That never-skilling occurs. The authors disclaim this themselves and do so more than once: 'Direct causal evidence linking AI exposure during training to competency failure in medical trainees does not exist', and the abstract concedes that direct evidence from medical training is absent. It is a risk model, explicitly not an established phenomenon. Prevalence, severity and reversibility are all stated as unknown, and the proposed framework is untested. Cite it for the taxonomy and as the origin of the terms, never as evidence of harm.
Working paper#
Liu, G., Christian, B., Dumbalska, T., Bakker, M. A. and Dubey, R. (2026). AI Assistance Reduces Persistence and Hurts Independent Performance#
arXiv:2604.04721, submitted 6 April 2026, revised 5 August 2026 (v4). Preprint, not peer reviewed
- Method
- A series of randomised controlled trials on human-AI interactions, N = 1,222, across mathematical reasoning and reading comprehension. AI assistance is available during a practice phase and then withdrawn, and performance is measured unassisted.
- Finding
- AI assistance improves performance in the short term, and people then perform significantly worse without AI and are more likely to give up. The authors report that these effects emerge after only brief interactions, approximately 10 minutes. They attribute the loss of persistence to AI conditioning people to expect immediate answers, denying them the experience of working through challenges on their own, and note that persistence is one of the strongest predictors of long-term learning.
- What it supports
- That a withdrawal effect can be produced causally, in a randomised design, and that it appears far faster than anyone had assumed. It also moves the mechanism from knowledge to persistence, which is a different and more portable claim.
- What it does not support
- Anything about sustained professional practice. These are short online tasks and the measured effect is a within-session carry-over rather than skill decay, so it cannot show whether the effect compounds, persists beyond the session or transfers to complex work. A preprint, not peer reviewed.
Working paper#
Sankaranarayanan, S. (2026). Mitigating "Epistemic Debt" in Generative AI-Scaffolded Novice Programming using Metacognitive Scripts#
arXiv 2602.20206, 22 February 2026, revised 31 March 2026
- Method
- Between-subjects experiment, 78 participants recruited via Prolific and UserInterviews.com, using a custom Cursor IDE plugin backed by Claude 3.5 Sonnet. Three conditions: manual control, unrestricted AI, and scaffolded AI. Followed by a 30-minute AI-blackout maintenance task. Preprint, not peer reviewed.
- Finding
- Both AI groups outperformed the manual control on functional utility (p < .001) and did not differ from each other (p = .64). On the subsequent AI-blackout maintenance task, unrestricted AI users failed at 77 per cent against 39 per cent for the scaffolded group. The author describes "fragile experts": developers whose high functional utility masks critically low corrective competence.
- What it supports
- That the gap between producing output and being able to repair it can be created inside a single session, and that interface design changes the size of that gap. The scaffolded condition roughly halved the failure rate.
- What it does not support
- A durable effect on capability. It is a single session with a single blackout task, in novice programming, with no longitudinal follow-up. The author puts "epistemic debt" in quotation marks in his own title and builds explicitly on Kirschner, so it is not offered as a new construct.
Working paper#
Shen, J. H. and Tamkin, A. (2026). How AI Impacts Skill Formation#
arXiv:2601.20245. Work conducted under the Anthropic Fellows Program
- Method
- Randomised experiments with developers learning a new asynchronous programming library through a self-guided tutorial, with and without an AI assistant.
- Finding
- AI use impaired conceptual understanding, code reading and debugging without significant average efficiency gains. Participants with AI assistance scored 17 per cent lower on a comprehension quiz than those who coded by hand, while finishing only marginally faster. The authors identify six AI interaction patterns, three of which involve cognitive engagement and preserve learning outcomes even with AI assistance.
- What it supports
- That the effect on learning depends on how the tool is used rather than on whether it is used, which is the same conclusion the guardrailed and scaffolded arms of the school and programming trials reached from the other direction.
- What it does not support
- That AI use degrades professional expertise. One library, one tutorial, developers rather than a general population, and a preprint. The conflict of interest runs against the finding rather than towards it, since the work was conducted under an AI company's fellowship and reports harm, but it is disclosed here either way.
Working paper#
Stromberg, D., Lei, V. and Wu, Y. (2026). The Generative AI Learning Penalty: Evidence from Chinese Secondary Education#
CEPR Discussion Paper 21577, CEPR Press, published 2 June 2026. Working paper, not peer reviewed
- Method
- Thirty months of panel data on 26,811 Chinese students in grades 7 to 12, combining monthly closed-book exams, high-school and college entrance exams, and homework scores and completion time across nine subjects. Staggered AI adoption in a difference-in-differences design. The outcome measures are closed-book and invigilated, so they record unaided performance by construction.
- Finding
- AI adoption raises homework scores by 18 per cent and reduces completion time by 30 per cent, and lowers monthly exam scores by 20 per cent within six months. High-stakes entrance-exam scores fall by 18 and 24 per cent, with the full penalty emerging only after about two years. Losses are largest in social science, then STEM, then languages, and are especially large for junior students, high-achieving students and boys. They concentrate among roughly 80 per cent of AI users whose behaviour is consistent with homework outsourcing, indicated by very short completion time coupled with high homework scores. Users who maintain similar completion time to non-users experience small losses.
- What it supports
- That the gap between assisted output and unaided capability can be measured at scale over years rather than minutes, and that it widens rather than closing. It also locates the damage in delegation rather than in access: the students who kept working at their normal pace were largely spared.
- What it does not support
- Causation with the confidence of a randomised trial. Adoption is self-selected and staggered rather than assigned, the outsourcing split is inferred from time-on-homework rather than observed, and a two-year lag makes contemporaneous confounds harder to exclude. One country, one school system, secondary students. Nothing about professional work. Not peer reviewed.
Human and AI collaboration
What actually happens when a person and a model work together.
Peer-reviewed#
Bainbridge, L. (1983). Ironies of Automation#
Automatica, 19(6)
- Method
- Theoretical analysis of automated process control systems and the human roles left within them.
- Finding
- Automating the routine parts of a task leaves the human with the hardest residue, monitoring and exception handling, while removing the routine practice that built the competence to do it. Automation makes the remaining human role harder, not easier.
- What it supports
- That monitoring is a demanding task rather than a light one, and that the design of automation determines whether the human retains the capability to supervise it.
- What it does not support
- Anything specific to AI. It is a process-control argument from 1983, and its application to generative systems is by analogy rather than by measurement.
Peer-reviewed#
Helmreich, R. L., Merritt, A. C. and Wilhelm, J. A. (1999). The Evolution of Crew Resource Management Training in Commercial Aviation#
International Journal of Aviation Psychology, 9(1)
- Method
- Historical and evaluative review of CRM programmes from 1979 onwards.
- Finding
- CRM began with a 1979 NASA workshop prompted by an NTSB finding that a captain had failed to accept input from junior crew. Line audits show CRM produces the intended behavioural change, though measured attitudes decay over time even with recurrent training.
- What it supports
- That an industry built a training response to a named human-factors failure and measured behaviour rather than opinion.
- What it does not support
- That CRM reduces accidents. The authors state directly that accident rates are too rare to serve as a validation criterion.
Peer-reviewed#
Dekker, S. W. A. and Woods, D. D. (2002). MABA-MABA or Abracadabra? Progress on Human-Automation Co-ordination#
Cognition, Technology & Work, 4(4), 240-244
- Method
- Conceptual analysis of function allocation methods in human factors. No data, no participants, no experiment.
- Finding
- Substitution-based function allocation, of which the Fitts list is the archetype, cannot deliver human-automation coordination, because the effects of automation are qualitative rather than quantitative. The authors name the underlying assumption the substitution myth and write that capitalising on a strength of automation does not replace a human weakness but creates new human strengths and weaknesses, often in unanticipated ways. Allocating a function also creates new functions for the other partner that did not exist before. They add that neither the list nor much of the supervisory control literature explains the cognitive work involved in deciding how and when to intervene or how to switch from level to level.
- What it supports
- That splitting the tasks is the wrong unit of design, and that the coordination at the boundary has been the acknowledged gap in this literature since 2002.
- What it does not support
- How large the effect is, or anything measurable at all. It is an argument. Its supporting accident examples are cited rather than analysed.
Peer-reviewed#
Dietvorst, B. J., Simmons, J. P. and Massey, C. (2015). Algorithm Aversion: People Erroneously Avoid Algorithms After Seeing Them Err#
Journal of Experimental Psychology: General, 144(1)
- Method
- Five experiments.
- Finding
- After seeing an algorithm err, people abandon it even when it demonstrably outperforms them.
- What it supports
- That trust in a model moves for reasons unrelated to its accuracy.
- What it does not support
- That this holds for conversational AI, which is far more recent and feels different to use.
Peer-reviewed#
Logg, J. M., Minson, J. A. and Moore, D. A. (2019). Algorithm Appreciation: People Prefer Algorithmic to Human Judgment#
Organizational Behavior and Human Decision Processes, 151, 90-103
- Method
- Six experiments on estimates and forecasts.
- Finding
- People often weight algorithmic advice MORE heavily than human advice. Domain experts are the notable exception.
- What it supports
- Together with Dietvorst, that miscalibration runs in both directions and cannot be fixed by telling people to use judgement.
- What it does not support
- Which tendency dominates in any given workplace.
Peer-reviewed#
Zhang, Y., Liao, Q. V. and Bellamy, R. K. E. (2020). Effect of confidence and explanation on accuracy and trust calibration in AI-assisted decision making#
Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency (FAT* 20)
- Method
- Two online experiments on an income-prediction task from US census data, 40 trials each. Experiment 1: 72 Mechanical Turk participants, nine per cell in a two by two by two design. Experiment 2: nine further participants, analysed against two Experiment 1 cells for a total of 27. Human unaided accuracy 65 per cent, model accuracy 75 per cent.
- Finding
- Showing confidence scores significantly increased trust, F(1,64)=4.64, p=.035, and significantly improved trust calibration when model confidence was above 80 per cent, F(4,256)=15.8, p<.001. There was no significant difference in AI-assisted accuracy across the prediction and confidence conditions, a result the authors report as rejecting their own hypothesis. Local explanations produced no significant change against baseline on switching behaviour or accuracy, with a reverse trend on accuracy.
- What it supports
- That a confidence score moves how a person feels about a system more reliably than it moves what they catch, and that explanation alone did neither in this setting.
- What it does not support
- Much on its own. Nine participants per cell in Experiment 1 and 27 in Experiment 2, no effect sizes, confidence intervals or standard deviations reported, no exclusions or attention checks described, non-expert participants, and a contrived task carrying no responsibility. The authors also note the approach depends on the model's probabilities being well calibrated in the first place, which the calibration literature says they usually are not.
Peer-reviewed#
Bucinca, Z., Malaya, M. B. and Gajos, K. Z. (2021). To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making#
Proceedings of the ACM on Human-Computer Interaction, 5(CSCW1), April 2021. DOI 10.1145/3449287
- Method
- Experiment with 199 participants comparing three cognitive forcing interventions, designed from dual-process theory to compel more thoughtful engagement with AI-generated explanations, against two simple explainable-AI approaches and a no-AI baseline. Includes an audit for intervention-generated inequalities using the Need for Cognition scale.
- Finding
- Cognitive forcing significantly reduced overreliance compared with the simple explainable-AI approaches. Participants gave the least favourable subjective ratings to the designs that reduced overreliance the most. On average the interventions benefited participants higher in Need for Cognition more, so human cognitive motivation moderates the effectiveness of explainable AI. The authors argue people rarely engage analytically with each individual recommendation and instead develop general heuristics about when to follow the AI.
- What it supports
- That deliberately adding friction to an AI-assisted decision measurably reduces acceptance of wrong suggestions, and that the intervention people dislike most is the one that works best.
- What it does not support
- That slower decisions produce better organisational outcomes. This is a controlled task with 199 participants, not a field study, and it measures overreliance rather than downstream results. The uneven benefit by Need for Cognition means the effect will not be uniform across a workforce.
Working paper#
Dell'Acqua, F. et al. (2023). Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality#
Harvard Business School and BCG working paper
- Method
- Field experiment, 758 BCG consultants, tasks inside and just outside GPT-4's competence.
- Finding
- Inside the frontier, AI-assisted consultants were dramatically better and faster. Outside it, they performed worse than consultants with no AI at all.
- What it supports
- That model competence is jagged rather than smooth, and that confident output suppresses scrutiny at exactly the wrong moment.
- What it does not support
- Where the frontier runs in your domain. That is local and must be learned.
Peer-reviewed#
Noy, S. and Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence#
Science, 381(6654), 187-192, 13 July 2023. DOI 10.1126/science.adh2586. Preregistered, AEA RCT Registry trial 10882
- Method
- Preregistered online randomised experiment. 453 college-educated professionals, including marketers, grant writers, consultants, data analysts, HR professionals and managers, given occupation-specific incentivised writing tasks of 20 to 30 minutes. Half were randomly exposed to ChatGPT. Run 27 January to 21 February 2023 with GPT-3.5.
- Finding
- Average time taken fell by 40 per cent and output quality rose by 18 per cent. Time on the post-treatment task dropped by 11 minutes, 0.75 standard deviations, against a control mean of 27 minutes, and evaluator grades rose by 0.45 standard deviations. Inequality between workers decreased: in the treatment group initial inequalities were more than half-erased, with the correlation between first-task and second-task grades falling to 0.14.
- What it supports
- That the tool compresses the performance distribution on tasks it does well, raising the floor much more than the ceiling. That is a pricing fact about expertise before it is a productivity fact.
- What it does not support
- That the effect generalises. The authors say they examined a limited range of occupations and tasks in which ChatGPT may be unusually useful, and speculate that real-economy effects will be somewhat lower. Short one-off tasks, online sample, an early model. Nothing about capability retention.
Working paper#
Peng, S., Kalliamvakou, E., Cihon, P. and Demirer, M. (2023). The Impact of AI on Developer Productivity: Evidence from GitHub Copilot#
arXiv:2302.06590, 13 February 2023. DOI 10.48550/arXiv.2302.06590. Authors at Microsoft Research, GitHub and MIT Sloan. Not peer reviewed
- Method
- Randomised controlled trial. 95 professional programmers recruited through Upwork from 166 offers, randomised to 45 treated and 50 control, given a standardised task of implementing an HTTP server in JavaScript. 35 in each group completed the task and survey.
- Finding
- Conditional on completion, the treated group averaged 71.17 minutes against 160.89 for control, a 55.8 per cent reduction in completion time, p = 0.0017, with a 95 per cent confidence interval on the improvement of 21 to 89 per cent. Participants in both groups estimated a 35 per cent productivity increase, which the authors describe as an underestimation of the 55.8 per cent revealed increase.
- What it supports
- A large measured speed gain on a well-specified task, and, separately, that self-reported productivity can UNDERSTATE the measured effect. Any claim that self-report systematically overstates gains has to answer this result.
- What it does not support
- Anything about quality. The authors state the study does not examine the effects of AI on code quality. The effect is estimated on 35 completers per arm rather than 95 participants, the task is standardised and greenfield rather than work inside a mature codebase, the authors are employed by the vendor and its parent. Not peer reviewed.
Peer-reviewed#
Zamfirescu-Pereira, J. D., Wong, R. Y., Hartmann, B. and Yang, Q. (2023). Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM Prompts#
CHI 2023
- Method
- Design probe study with non-experts using a purpose-built prompt design tool.
- Finding
- Non-experts approached prompting opportunistically rather than systematically, over-generalised from single successes and failures, and struggled to form an accurate model of the system.
- What it supports
- That prompting is genuinely harder than it looks, which is the strongest case FOR teaching it.
- What it does not support
- That this persists. It used 2023 models, and providers are actively engineering the difficulty away.
Peer-reviewed#
Kim, S. S. Y., Liao, Q. V., Vorvoreanu, M., Ballard, S. and Wortman Vaughan, J. (2024). I'm Not Sure, But...: Examining the Impact of Large Language Models' Uncertainty Expression on User Reliance and Trust#
Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (FAccT 24). Pre-registered at osf.io/mnrp9
- Method
- Pre-registered between-subjects experiment, four conditions, eight yes-or-no medical questions. 656 responses collected, 252 excluded on pre-registered criteria, final sample 404. Responses pre-generated and presented as a fictional system whose answers were correct on exactly half the questions. Conditions varied only the presence and perspective of uncertainty expression.
- Finding
- Access to the system raised agreement to 80.9 per cent against 58.4 per cent without, and lowered accuracy to 63.9 per cent against 74.2 per cent without. First-person uncertainty expression significantly reduced agreement to 74.8 per cent and significantly raised accuracy to 72.8 per cent. Impersonal expression moved both in the same direction without reaching significance. Intention to use fell significantly under first-person hedging, 2.91 against 3.25 in control and 3.36 for the impersonal version. Hedging reduced accuracy slightly when the system was correct and raised it more when the system was wrong.
- What it supports
- That hedging changes reliance in the direction that helps, that the perspective of the hedge matters, and that the version which helps most is the version users least want to keep using.
- What it does not support
- That uncertainty expression should be mandated. The authors state directly that regulators should avoid blanket requirements until more research is done, and flag that their system had low accuracy and expressed uncertainty often in a poorly calibrated manner. It did not eliminate over-reliance: participants without AI access still performed best. Reported as model-estimated means with significance stars, no test statistics, exact p values or effect sizes, and 38.4 per cent of collected responses were excluded.
Peer-reviewed#
Vaccaro, M., Almaatouq, A. and Malone, T. (2024). When combinations of humans and AI are useful: a systematic review and meta-analysis#
Nature Human Behaviour, 8, 2293-2303
- Method
- Preregistered systematic review and meta-analysis: 106 experimental studies, 370 effect sizes, published January 2020 to June 2023.
- Finding
- Human-AI combinations performed significantly WORSE on average than the better of human or AI alone (Hedges' g = -0.23). Losses concentrated in decision-making; gains in content creation. Pairing gained where humans beat the AI and lost where the AI beat humans.
- What it supports
- That adding a human is not a control, and that undesigned pairing can subtract. The most under-absorbed result in the field.
- What it does not support
- That human-AI teams are useless. The benchmark is an oracle-selected best performer, which you rarely know in advance. Also predates current frontier models.
Peer-reviewed#
Xiong, M., Hu, Z., Lu, X., Li, Y., Fu, J., He, J. and Hooi, B. (2024). Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs#
ICLR 2024. arXiv:2306.13063
- Method
- Systematic evaluation of confidence elicitation across five models, Vicuna 13B, GPT-3, GPT-3.5-turbo, GPT-4 and LLaMA 2 70B, on eight datasets spanning arithmetic, commonsense, symbolic and professional reasoning. Metrics: expected calibration error, AUROC and AUPRC. No human subjects.
- Finding
- Average expected calibration error for plain verbalised confidence is 0.520 for GPT-3, 0.461 for Vicuna, 0.436 for LLaMA 2, 0.377 for GPT-3.5 and 0.180 for GPT-4. GPT-4's average AUROC is 62.7 per cent against a 50 per cent chance baseline. Stated confidences cluster in the 80 to 100 per cent range in multiples of five, which the authors suggest means models may be imitating human expressions when verbalising confidence.
- What it supports
- That a number a model gives when asked how sure it is is close to unusable as a calibrated quantity, and that its shape suggests it is generated as plausible text rather than measured.
- What it does not support
- Anything about how users read these numbers, since no humans were involved. Mitigation strategies reduce ECE substantially, to 0.028 in the best case, while still failing to predict incorrect answers on knowledge-heavy tasks. Models are of the GPT-4 and LLaMA 2 generation.
Peer-reviewed#
Zhou, K., Hwang, J. D., Ren, X. and Sap, M. (2024). Relying on the Unreliable: The Impact of Language Models' Reluctance to Express Uncertainty#
Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024), Long Papers, 3623-3643. Also arXiv:2401.06730
- Method
- Model study: nine models prompted with 49 prompts over 284 MMLU questions, 125,244 queries. Human study: Prolific participants across a control setting and three interactive settings with calibrated, overconfident and underconfident systems, 25 recruited per setting.
- Finding
- Only about 5 per cent of generated answers include any epistemic marker. Among confidently expressed responses the error rate averages 47 per cent, and only 53 per cent of generations expressing certainty are correct. In the human study, hedged answers were relied on around 10 per cent of the time and confident ones around 90, but plain unmarked statements were also relied on nearly 90 per cent of the time, which the authors read as users interpreting the absence of a marker as certainty. Exposure to an overconfident system left mental models uncorrected, with participants averaging 76 per cent during miscalibrated rounds and 86 afterwards, while an underconfident system produced 66 and then 98. Reward modelling scores plain statements at 4.03, expressions of certainty at 0.82 and expressions of doubt at minus 1.86.
- What it supports
- That silence about uncertainty is read as confidence rather than as neutrality, and that the confident register is a product of a preference model that penalises hedging more than it rewards assurance.
- What it does not support
- Any of it with statistical rigour on the human side. No total human N is stated in the text, and no p values, confidence intervals or effect sizes are reported for any human result. Twenty-five participants per setting. US-only, which the authors themselves call a narrow and US-centric view. Body text and Table 1 disagree slightly on the marker rate, 5 against 6 per cent. Peer-reviewed venue corrected to ACL 2024 on 4 September 2026 after an earlier draft of this entry read the arXiv preprint as unpublished.
Peer-reviewed#
Cemri, M., Pan, M. Z., Yang, S., Agrawal, L. A., Chopra, B., Tiwari, R., Keutzer, K., Parameswaran, A., Klein, D., Ramchandran, K., Zaharia, M., Gonzalez, J. E. and Stoica, I. (2025). Why Do Multi-Agent LLM Systems Fail?#
NeurIPS 2025 Datasets and Benchmarks Track. arXiv:2503.13657
- Method
- Empirical failure taxonomy built from expert annotation of 150 execution traces at inter-annotator kappa 0.88, then applied at scale with LLM-as-judge across more than 1,600 traces from seven multi-agent frameworks. Failure distribution computed on 210 traces.
- Finding
- Failures divide into system design issues 41.8 per cent, inter-agent misalignment 36.9 per cent and task verification 21.3 per cent. Inter-agent misalignment is defined as a breakdown in critical information flow during interaction and coordination, and comprises conversation resets 2.20 per cent, proceeding on wrong assumptions rather than seeking clarification 6.80 per cent, task derailment 7.40 per cent, withholding information another agent needed 0.85 per cent, ignoring another agent's input 1.90 per cent, and mismatch between reasoning and action 13.2 per cent. The authors state that context and communication protocols are often insufficient, because the errors occur even when agents in the same framework communicate in natural language. Improving role specification alone raised ChatDev success by 9.4 percentage points.
- What it supports
- That more than a third of observed multi-step agent failure sits at the joins rather than in any single agent's competence, and that standardising the message format does not fix it.
- What it does not support
- Anything about human-agent boundaries. Every handoff in the dataset is agent to agent and no humans appear anywhere. The authors caveat that the 210-trace distribution illustrates system-specific profiles rather than comparing performance across systems, so the percentages are not cross-benchmark comparable.
Peer-reviewed#
Dell'Acqua, F., Ayoubi, C., Lifshitz, H., Sadun, R., Mollick, E., Mollick, L., Han, Y., Goldman, J., Nair, H., Taub, S. and Lakhani, K. (2025). The Cybernetic Teammate: A Field Experiment on Generative AI Reshaping Teamwork and Expertise#
NBER Working Paper 33641, April 2025. Published as 'The Cybernetic Teammate: A Field Experiment on Generative AI and Teamwork', Organization Science, June 2026. DOI 10.1287/orsc.2025.20702
- Method
- Pre-registered field experiment with 776 professionals at Procter and Gamble working on real product innovation challenges, randomised both on AI access and on working individually or in a two-person new product development team.
- Finding
- Individuals working with AI matched the performance of two-person teams working without it. AI use removed the functional split in proposals: without AI, research and development professionals proposed more technical solutions and commercial professionals more commercially oriented ones, while professionals using AI produced balanced solutions regardless of background. Participants using AI reported more positive emotional responses.
- What it supports
- That a model can substitute for measurable parts of what a second human teammate contributes, including some of the social and motivational function, and that it flattens the differences in output that come from professional training.
- What it does not support
- Client or business outcomes, which were not measured, or any effect over time. One firm, one task type, a single session. Procter and Gamble provided financial support to the institute involved and one author had consulted for the firm, both disclosed in the paper. That the flattened output is better rather than merely more balanced is not established.
Working paper#
Model Evaluation and Threat Research (METR) (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity#
METR, July 2025
- Method
- Randomised controlled trial. 16 experienced open-source developers, 246 real tasks on mature repositories, each task randomly assigned to permit or prohibit AI tools.
- Finding
- Developers were measured as 19 percent SLOWER when permitted to use AI tools. They had forecast a 24 percent speed-up beforehand, and after completing the tasks and experiencing the slowdown, still estimated AI had made them about 20 percent faster.
- What it supports
- That self-reported productivity gain is an unreliable measure of actual productivity gain, and that the error can run in the opposite direction to the truth by a wide margin.
- What it does not support
- That AI slows all developers or all software work, and NOT the current position. Sixteen participants, all experienced, all working on large mature codebases they knew well, using early-2025 tooling. METR themselves withdrew this as a current signal on 24 February 2026: see metr-2026-update. The durable finding is the perception gap, not the 19 per cent.
Peer-reviewed#
Steyvers, M., Tejeda, H., Kumar, A., Belem, C., Karny, S., Hu, X., Mayer, L. W. and Smyth, P. (2025). What large language models know and what people think they know#
Nature Machine Intelligence, 7, 221-231
- Method
- Two behavioural experiments, 301 participants recruited through Prolific, each assigned 40 questions from pools of 350 multiple-choice and 336 short-answer items. Model confidence from GPT-3.5, PaLM2 and GPT-4o compared against participant confidence after reading model explanations.
- Finding
- Model confidence discriminates correct from incorrect answers at AUC 0.751 for GPT-3.5, 0.746 for PaLM2 and 0.781 for GPT-4o. Participants reading default explanations reached 0.589, 0.602 and 0.592, which the authors describe as only slightly better than random guessing. They name the shortfalls the calibration gap and the discrimination gap, and attribute human miscalibration primarily to overconfidence, people believing LLMs are more accurate than they are. Longer explanations significantly raised participant confidence without improving discrimination, mean participant AUC 0.54 for long explanations.
- What it supports
- That the reader is a worse judge of an answer's correctness than the model is, on the same items, and that adding explanation length makes readers surer rather than righter.
- What it does not support
- That closing the gap improves task accuracy. The modified-explanation result is a simulation via post-hoc filtering rather than a live deployment. Participants had no domain expertise and their own accuracy was 33 per cent against the model's 39. Statistics are Bayes factors only, with no p values or effect sizes, and the ECE values sit inside a figure rather than in the text.
Working paper#
Becker, J. (METR) (2026). Measuring the Self-Reported Impact of Early-2026 AI on Technical Worker Productivity#
METR, 11 May 2026. Not peer reviewed
- Method
- Survey of 349 technical workers on self-reported change in speed and in value of work attributable to AI tools, set against the published field-experiment literature.
- Finding
- Median self-reported change in the VALUE of work is 1.4 to 2 times, and median self-reported SPEED change is 3 times. METR state that to their knowledge only one study has gathered survey and field experiment results on the same population and metric, Becker et al. (2025), which finds developers overestimate productivity gains by over 40 percentage points. They add that public survey estimates have tended to exceed field-experiment estimates, while stating it is difficult to determine the extent to which surveys overestimate gains relative to experimental data, and that speed measures are likely biased upwards relative to value measures.
- What it supports
- That the evidence base for the most repeated claim in enterprise AI, that surveys overstate productivity, rests on a single population, and that the people who own that finding say so themselves. It also separates speed from value, which almost no business case does.
- What it does not support
- That self-report always overstates. The GitHub Copilot randomised trial found the opposite direction, with participants estimating 35 per cent against a measured 55.8 per cent. A survey of a self-selected technical population, not peer reviewed, and its own authors decline the general claim.
Working paper#
Becker, J., Rush, N., Cunningham, T., Rein, D. and Mahamud, K. (METR) (2026). We are Changing our Developer Productivity Experiment Design#
METR, 24 February 2026
- Method
- Second randomised task-level study, begun August 2025: 57 developers (10 returning from the original study, 47 newly recruited), 143 repositories, more than 800 tasks, paid 50 dollars an hour against 150 in the original. Accompanied by participant surveys and interviews.
- Finding
- METR state the data gives an unreliable signal of the current productivity effect of AI tools, because 30 to 50 per cent of developers reported declining to submit tasks they did not want to do without AI, and an increased share declined to take part at all. Raw results now point the other way: an estimated speedup of -18 per cent for returning developers (CI -38 to +9) and -4 per cent for new recruits (CI -15 to +9), against the original +19 per cent slowdown (CI +2 to +39). They believe developers are likely more sped up in early 2026 than in early 2025, while stating their own data is only very weak evidence for the size of that change.
- What it supports
- That the 19 per cent slowdown belongs to early 2025 and should not be quoted as the current effect. It also demonstrates a measurement problem that will worsen: as adoption rises, the people most helped by AI are the ones most likely to select themselves out of any study that asks them to work without it.
- What it does not support
- That AI now speeds developers up by a specific amount. Every confidence interval here crosses zero, and METR say so. It does not retract the original study, whose perception-gap finding, that participants estimated a 20 per cent speed-up while measured slower, is untouched.
Working paper#
Tomasev, N., Franklin, M. and Osindero, S. (2026). Intelligent AI Delegation#
Google DeepMind. arXiv:2602.11865, submitted 12 February 2026. Preprint, not peer-reviewed
- Method
- A framework paper proposing how AI agents should decompose problems and delegate across other agents and people, covering task assignment, monitoring, trust, permissions and accountability in long delegation chains. No experiment, no data. The capability material is section 5.6, Risk of De-skilling.
- Finding
- The authors name oversight readiness as a property a future workforce may lack, arguing that expertise is built through the repetitive execution of narrowly scoped tasks, that those are the tasks most likely to be delegated to agents first, and that fully automating them would deprive junior staff of the experience needed for strategic judgement. Their proposed remedies are unusual for a frontier lab: curriculum-aware task routing that allocates work inside a junior's zone of proximal development, and a delegation framework that should occasionally introduce minor inefficiencies by routing tasks to humans deliberately, to maintain their skills.
- What it supports
- That the missing rungs argument is being reached independently by people building delegation infrastructure rather than only by those writing about its consequences, and that at least one frontier lab has proposed deliberate inefficiency as a capability-preservation mechanism.
- What it does not support
- Anything measured. It is a preprint framework paper with no empirical component, so it establishes that the risk is taken seriously by practitioners and not that the risk has been observed. The proposed routing systems also assume an organisation can assess what a junior can currently do, which is the unsolved part.
Empathy, creativity and what stays human
Where machine performance already exceeds ours, and what that does not settle.
Peer-reviewed#
Ayers, J. W. et al. (2023). Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum#
JAMA Internal Medicine, 183(6), 589-596
- Method
- Cross-sectional study, 195 real patient questions from a public forum, blind-rated by licensed healthcare professionals.
- Finding
- Chatbot responses were rated good or very good quality 78.5 percent of the time against 22.1 percent for physicians, and empathetic or very empathetic 45.1 percent against 4.6 percent.
- What it supports
- That on the observable, textual performance of empathy, the machine already wins comfortably. The claim that empathy is safe from AI is empirically wrong.
- What it does not support
- That a machine can care for anyone. Doctors answering strangers free of charge between patients are not doing the job they trained for.
Peer-reviewed#
Hohenstein, J., Kizilcec, R. F., DiFranzo, D., Aghajari, Z., Mieczkowski, H., Levy, K., Naaman, M., Hancock, J. and Jung, M. F. (2023). Artificial intelligence in communication impacts language and social relationships#
Scientific Reports, 13, 5487. DOI 10.1038/s41598-023-30938-9
- Method
- Two randomised experiments on algorithmic response suggestions (smart replies) in live text chat. Study 1: 438 Mechanical Turk crowdworkers in 219 pairs discussing a policy question, smart-reply availability randomised separately for each partner, analysed with an instrumental-variable design. Study 2: 582 crowdworkers in 291 pairs, with the sentiment of the suggested replies manipulated. Preregistered (AsPredicted #40389).
- Finding
- Smart replies accounted for 14.3 per cent of messages and produced 10.2 per cent more messages per minute. Greater ACTUAL use by a partner improved the other person's rating of their cooperation (b=15.66, p=0.018) and felt affiliation towards them (b=21.79, p=0.007), with no effect on dominance. Greater SUSPECTED use had the opposite sign: the more a participant believed their partner used smart replies, the less cooperative (p<0.0001) and less affiliative (p<0.0001) they rated them, after controlling for actual use. Suspicion tracked actual use only weakly (Pearson's r=0.22).
- What it supports
- That the interpersonal penalty for AI-assisted messaging attaches to being suspected rather than to using it, and that suspicion is a poor detector. Actual use made people seem warmer, not colder.
- What it does not support
- Anything longitudinal, which the authors state directly. The suspicion result is correlational and they say it does not show causally how attitudes shift in response to actual use. Participants were crowdworkers discussing policy with strangers, not intimates. A publisher correction (10.1038/s41598-023-43601-0, 3 October 2023) added an omitted funding statement and changed no result.
Peer-reviewed#
Jakesch, M., Bhat, A., Buschek, D., Zalmanson, L. and Naaman, M. (2023). Co-Writing with Opinionated Language Models Affects Users' Views#
CHI 2023, ACM
- Method
- Online experiment with 1,506 participants writing a post on whether social media is good for society, assisted by a tool configured to argue for or against. Opinions rated by 500 independent judges, plus a post-task attitude survey.
- Finding
- The opinionated model changed both the opinions expressed in participants' writing and their own opinions in the subsequent attitude survey. The effect held among participants who had ample time to write independently. The authors call it latent persuasion.
- What it supports
- That writing assistance moves what people think, not only what they type.
- What it does not support
- Generalisation across topics, or persistence after the task. One topic, one configuration, self-reported attitudes.
Peer-reviewed#
Doshi, A. R. and Hauser, O. P. (2024). Generative AI enhances individual creativity but reduces the collective diversity of novel content#
Science Advances, 10(28)
- Method
- Online experiment, 293 writers producing short fiction and 600 evaluators.
- Finding
- AI-assisted stories were rated more creative, better written and more enjoyable, with the largest gains for the least creative writers, and were markedly more similar to one another.
- What it supports
- That individual creative quality and collective creative range move in opposite directions. A social dilemma: every writer is right to use it, and the literature gets duller.
- What it does not support
- That this generalises beyond one short creative task with one form of assistance.
Peer-reviewed#
Yin, Y., Jia, N. and Wakslak, C. J. (2024). AI can help people feel heard, but an AI label diminishes this impact#
PNAS, 121(14), e2319112121
- Method
- Experiments comparing AI-generated and human-written responses, with and without disclosure.
- Finding
- AI-generated replies made recipients feel MORE heard than replies from untrained humans, and labelling the reply as AI removed the advantage.
- What it supports
- That the value of recognition is not in the words but in the belief that a person chose to attend to you. The most clarifying study in this debate.
- What it does not support
- That the label effect is stable. Norms around disclosed AI assistance are moving, and nobody has measured this over time.
Institutional survey#
Common Sense Media (2025). Talk, Trust, and Trade-offs: How and Why Teens Use AI Companions#
Common Sense Media, San Francisco, 16 July 2025
- Method
- Nationally representative survey of 1,060 US teens aged 13 to 17, April and May 2025. Risk items were asked of the 758 respondents who had used an AI companion.
- Finding
- 72 per cent of teens had used an AI companion at least once and 52 per cent were regular users, a few times a month or more. Among users, 33 per cent had chosen an AI companion over a real person for something important or serious, 34 per cent had felt uncomfortable with something a companion said or did, and 24 per cent had shared personal or private information. The same survey found 67 per cent of teens rating AI conversations as less satisfying than conversations with real friends against 10 per cent more satisfying, 80 per cent of users spending more time with friends than with companions, and 50 per cent not trusting the advice.
- What it supports
- That AI companion use among US teenagers is majority behaviour rather than a fringe one, and that a substantial minority of users have taken serious conversations and personal information to them.
- What it does not support
- Anything about developmental effects, which are not measured here. Self-report at a single point in time, by an organisation that backs legislation to ban these products for minors, and the risk percentages are of users rather than of all teens: 33 per cent of users choosing a companion for a serious conversation is roughly 24 per cent of teens. The press release framing of a generation replacing human connection is not supported by the survey's own Q5 and Q7.
Peer-reviewed#
Joshi, N. and Vogel, D. (2025). Writing with AI Lowers Psychological Ownership, but Longer Prompts Can Help#
ACM Conversational User Interfaces 2025
- Method
- Two within-subjects experiments, 31 and 34 participants, writing short stories across conditions from a three-word prompt to writing unaided.
- Finding
- Psychological ownership rose steadily with prompt length, from a mean of 1.80 with a three-word prompt to 6.29 writing alone. The benefit plateaued once the prompt reached roughly the length of the target text, and no AI-assisted condition reached the ownership of writing unaided.
- What it supports
- That how much of yourself you put in changes whether the output feels like yours, and by a large margin.
- What it does not support
- Anything about professional or long-form writing. Short fiction, small samples.
Compiled review#
Sourati, Z., Ziabari, A. S. and Dehghani, M. (2026). The Homogenizing Effect of Large Language Models on Human Expression and Thought#
Trends in Cognitive Sciences
- Method
- Synthesis across linguistics, psychology, cognitive science and computer science. Not an original experiment.
- Finding
- Argues that models reflect and reinforce dominant styles while marginalising alternatives, and that reliance on a small number of systems amplifies convergence across users.
- What it supports
- That the concern is taken seriously across several fields rather than being a commentator's intuition.
- What it does not support
- Any specific stylistic feature converging. It presents no new measurement of its own.
Work, jobs and the labour market
Exposure, adoption and measured effect, which are three different things.
Compiled review#
Barney, J. (1991). Firm Resources and Sustained Competitive Advantage#
Journal of Management, 17(1), 99-120, March 1991. DOI 10.1177/014920639101700108
- Method
- Theoretical article in strategic management. No new data. Builds on the stated assumptions that strategic resources are heterogeneously distributed across firms and that those differences are stable over time.
- Finding
- Sets out four empirical indicators of the potential of a firm resource to generate sustained competitive advantage: value, rareness, imitability and substitutability. Applies the model to several firm resources and draws out implications for other business disciplines. The founding statement of the resource-based view.
- What it supports
- That a widely used and long-established test exists for whether a resource can confer advantage, against which a purchasable AI licence can be assessed.
- What it does not support
- Anything empirical, and nothing about AI, which postdates it by three decades. It is a theory with a substantial critical literature, and applying it to a technology three years into commercial diffusion is an argument rather than a finding. The abstract read at source uses rareness, imitability and substitutability; the VRIN acronym is later shorthand and does not appear in the paper's abstract.
Peer-reviewed#
Garicano, L. (2000). Hierarchies and the Organization of Knowledge in Production#
Journal of Political Economy, 108(5), 874-904. University of Chicago Press
- Method
- Theoretical model of knowledge acquisition and problem-solving in production, with communication costs and knowledge acquisition costs traded off against each other. No empirical estimation.
- Finding
- A knowledge-based hierarchy is a natural way to organise the acquisition of knowledge when matching problems with those who know how to solve them is costly. Production workers acquire knowledge of the most common or easiest problems and refer exceptions upward to specialist problem solvers, with problems passed on until somebody solves them or the conditional probability of a solution is too low to justify continuing. Adding layers of problem solvers raises the utilisation rate of knowledge and economises on knowledge acquisition, at the cost of increasing the communication required.
- What it supports
- That the number of layers in an organisation is a function of two costs rather than of custom, which gives a testable prediction for what happens when either cost falls. It is the model every subsequent claim about AI flattening organisations implicitly relies on.
- What it does not support
- Anything measured. It is a model, published a quarter of a century before generative AI, and it contains no data about firms, layers or technology adoption.
Peer-reviewed#
Gathmann, C. and Schoenberg, U. (2010). How General Is Human Capital? A Task-Based Approach#
Journal of Labor Economics, 28(1)
- Method
- German administrative employment panel, using task overlap between occupations to measure task-specific human capital.
- Finding
- Task-specific human capital accounts for up to 52 per cent of overall wage growth. Workers move to occupations with similar task profiles, and the distance of those moves shrinks with experience.
- What it supports
- That skill is portable along task lines rather than being either fully general or locked to one job.
- What it does not support
- That broad general education transfers well. If anything it argues the opposite, which complicates the study-anything advice rather than supporting it.
Working paper#
Frey, C. B. and Osborne, M. A. (2013). The Future of Employment: How Susceptible Are Jobs to Computerisation?#
Oxford Martin School working paper, 17 September 2013. Later published in Technological Forecasting and Social Change, 114 (2017)
- Method
- Probability of computerisation estimated for 702 detailed US occupations using a Gaussian process classifier, trained on 70 occupations hand-labelled by machine learning researchers at an Oxford workshop.
- Finding
- In the authors' own words: 'about 47 percent of total US employment is at risk'. The paper's title asks how SUSCEPTIBLE jobs are, and the estimate is of technical susceptibility to computerisation, not a forecast of job losses.
- What it supports
- That a large share of US employment sits in occupations whose tasks were, in 2013, judged technically susceptible to computerisation.
- What it does not support
- That 47 per cent of jobs will be, or have been, lost. It is not a prediction, carries no date attached to any loss, and models whole occupations rather than tasks within them. The routine citation as '47 per cent of jobs will disappear' reverses what the paper claims. The working paper is 2013; the journal version is 2017, and the two dates are frequently confused.
Peer-reviewed#
Bloom, N., Garicano, L., Sadun, R. and Van Reenen, J. (2014). The Distinct Effects of Information Technology and Communication Technology on Firm Organization#
Management Science, 60(12), 2859-2885. DOI 10.1287/mnsc.2014.2013. Earlier version NBER Working Paper 14975, May 2009
- Method
- Survey data on worker and plant manager autonomy and span of control, combined with measures of technology adoption, and instrumented using distance from ERP's place of origin and heterogeneous telecommunication costs arising from regulation. The working-paper version describes approximately 1,000 manufacturing firms with 100 to 5,000 employees across the US, France, Germany, Italy, Poland, Portugal, Sweden and the UK, drawn from the CEP double-blind management survey, with technology data from the Harte-Hanks ICT panel.
- Finding
- Information technology is a decentralising force and communication technology is a centralising force. Better information technologies, ERP for plant managers and computer-assisted design or manufacturing for production workers, are associated with more autonomy and a wider span of control. Technologies that improve communication, such as data intranets, decrease autonomy for workers and plant managers. Instrumenting strengthens the result.
- What it supports
- That 'technology' has no single organisational effect, and that predicting what a tool does to authority requires knowing whether it lowers the cost of knowing or the cost of telling.
- What it does not support
- Anything about generative AI, which is both kinds of technology in one interface. Manufacturing plants only, data predating 2009. The sample description above is verified in the working paper; the published article abstract says only 'American and European manufacturing firms', and the typeset article could not be opened at the publisher.
Institutional modelling#
United States Bureau of Labor Statistics (2018). Occupational Projections Evaluation, 2006 to 2016#
BLS Employment Projections programme
- Method
- The agency's own retrospective scoring of its 2006 ten-year projections for 840 detailed occupations against actual 2016 outcomes.
- Finding
- BLS correctly projected whether an occupation would grow or decline 78 per cent of the time, but correctly projected which occupations would grow faster than the economy as a whole only 57 per cent of the time. Projected average occupational growth was 10.4 per cent against an actual 3.6 per cent, the gap driven by a recession the projections could not foresee.
- What it supports
- That the only organisation which scores its own occupational forecasts gets the useful question, relative growth, barely better than a coin toss, and that its largest errors come from shocks.
- What it does not support
- That forecasting is worthless. Direction of change at aggregate level was reasonably good. It also predates generative AI, and no equivalent scored record exists for a technological discontinuity.
Peer-reviewed#
Acemoglu, D. and Restrepo, P. (2019). Automation and New Tasks: How Technology Displaces and Reinstates Labor#
Journal of Economic Perspectives, 33(2), 3-30. DOI 10.1257/jep.33.2.3
- Method
- Task-based theoretical framework in which production is allocated between capital and labour, with an empirical decomposition of US industry-level data over recent decades.
- Finding
- Automation shifts the task content of production against labour through a displacement effect, and therefore ALWAYS reduces the labour share in value added, and may reduce labour demand even while raising productivity. The counterweight is the creation of new tasks in which labour has a comparative advantage, which always raises the labour share. Their decomposition attributes slower employment growth over three decades to an accelerating displacement effect, a weaker reinstatement effect and slower productivity growth.
- What it supports
- That the distributional question is separate from the productivity question, and that a technology can raise output while reducing labour's share of it. This is the framework almost every serious argument in this area now runs through.
- What it does not support
- That AI specifically will behave this way. The empirical work predates generative AI and concerns industrial automation and robotics.
Peer-reviewed#
Brynjolfsson, E., Rock, D. and Syverson, C. (2021). The Productivity J-Curve: How Intangibles Complement General Purpose Technologies#
American Economic Journal: Macroeconomics, 13(1), 333-72, January 2021. DOI 10.1257/mac.20180386. Earlier version NBER Working Paper 25148
- Method
- Theoretical model of general purpose technology adoption with unmeasured intangible complementary investment, applied to US national accounts data on computer hardware and software.
- Finding
- General purpose technologies enable and require significant complementary investments that are often intangible and poorly measured in national accounts. This produces underestimation of productivity growth in a new technology's early years and overestimation later, when the benefits of the intangible investments are harvested, a pattern the authors name the Productivity J-curve. Adjusting for intangibles related to computer hardware and software yields a total factor productivity level 15.9 per cent higher than official measures by the end of 2017. The authors state the AI-related intangible capital effects on measured productivity are currently small but growing.
- What it supports
- That an early productivity read on a general purpose technology is biased downwards for a structural and quantified reason, and that the direction of the later bias is upwards. It supplies the argument for staging an AI investment review rather than taking a single reading.
- What it does not support
- That AI will follow the same curve, or on what timescale. The 15.9 per cent figure is for computer hardware and software to 2017, not for AI, and the authors describe current AI intangible effects as small. It is also a measurement argument rather than a forecast of returns to any individual firm.
Peer-reviewed#
Yang, L., Holtz, D., Jaffe, S., Suri, S., Sinha, S., Weston, J., Joyce, C., Shah, N., Sherman, K., Hecht, B. and Teevan, J. (2022). The effects of remote work on collaboration among information workers#
Nature Human Behaviour, 6, 43-54. DOI 10.1038/s41562-021-01196-4, published online 9 September 2021
- Method
- Observed telemetry on emails, calendars, instant messages, video and audio calls and workweek hours of 61,182 US Microsoft employees over the first six months of 2020, using workers already remote before the pandemic as a comparison to separate firm-wide remote work from other pandemic effects.
- Finding
- Firm-wide remote work caused the collaboration network to become more static and siloed, with fewer bridges between disparate parts of the organisation, a decrease in synchronous and an increase in asynchronous communication. The authors state these effects may make it harder for employees to acquire and share new information across the network.
- What it supports
- That changing how information moves reorganises who knows what, without anybody redesigning a role, and that the change is measurable in the network rather than only in self-report.
- What it does not support
- Anything about AI, and nothing about outcomes. It measures communication structure rather than performance, in one very large technology company, during a pandemic. Full text is paywalled; this entry rests on the published abstract.
Peer-reviewed#
Acemoglu, D. (2024). The Simple Macroeconomics of AI#
NBER Working Paper 32487; published in Economic Policy, 40(121), 2025
- Method
- Task-based macroeconomic model applying Hulten's theorem to existing estimates of AI task exposure and task-level cost savings.
- Finding
- Estimates total factor productivity gains of no more than 0.66 percent over ten years, revised to under 0.53 percent once the difficulty of hard-to-learn tasks is accounted for. Argues AI is likely to widen the gap between capital and labour income rather than reduce labour income inequality.
- What it supports
- That plausible macroeconomic gains are an order of magnitude smaller than the headline value estimates in circulation.
- What it does not support
- That AI is unimportant. It models productivity through task-level cost savings, and would not capture effects running through new products, new tasks or capability change.
Working paper#
Autor, D. (2024). Applying AI to Rebuild Middle Class Jobs#
NBER Working Paper 32140
- Method
- Argument and synthesis rather than empirical test, drawing on the author's prior work on task structure and labour demand.
- Finding
- Argues that AI's distinctive opportunity is to extend the reach of expertise, letting a wider set of workers with complementary knowledge perform higher-stakes decision tasks currently reserved to elite experts.
- What it supports
- That there is a serious, well-argued case for AI as an expertise-widening technology rather than an expertise-replacing one.
- What it does not support
- That this will happen. The author is explicit that the thesis is an argument about what is possible rather than a forecast, and no evidence yet shows it occurring at scale.
Working paper#
Bick, A., Blandin, A. and Deming, D. J. (2024). The Rapid Adoption of Generative AI#
NBER Working Paper 32966
- Method
- Nationally representative US surveys of generative AI use at work and at home.
- Finding
- By late 2024, nearly 40 percent of US adults aged 18-64 used generative AI and 23 percent of employed respondents had used it for work in the previous week, but only 1 to 5 percent of all work hours were assisted.
- What it supports
- Enormous reach, thin penetration into actual hours. Adoption is not transformation.
- What it does not support
- Quality of use. Self-reported use counts any use at all.
Working paper#
Bonney, K., Breaux, C., Buffington, C., Dinlersoz, E., Foster, L., Goldschlag, N., Haltiwanger, J., Kroff, Z. and Savage, K. (2024). Tracking Firm Use of AI in Real Time: A Snapshot from the Business Trends and Outlook Survey#
US Census Bureau, Center for Economic Studies Working Paper CES 24-16, March 2024; also NBER Working Paper 32319. Not peer reviewed
- Method
- Analysis of the AI questions in the Business Trends and Outlook Survey, a high-frequency nationally representative survey of US firms, over the collection period covered by the paper.
- Finding
- Bi-weekly estimates of the AI use rate rose from 3.7 to 5.4 per cent, with an expected rate of about 6.6 per cent by early autumn 2024. 94.6 per cent of AI-using businesses reported no net change in employment in the previous six months attributable to AI use. On discontinuation, which the survey does not ask about directly, 67.9 per cent of current AI users expect to use it in future, 14.5 per cent do not, and 17.6 per cent do not know, so about one in seven current users may de-adopt; the authors attribute this to experimentation that does not yield anticipated benefits or organisational synergies. The authors state explicitly that the analysis does not seek to identify a causal link between AI use and firm performance, only whether use is associated with better performance in general, and that causal analysis awaits integrated data and repeat collections after about three to four years.
- What it supports
- That the body with the best firm-level data in the world declines to make the causal claim that consultancy and vendor reporting makes routinely, and says in its own text what would be required to make it.
- What it does not support
- Any effect of AI on firm performance, by the authors' own statement. Adoption rates here are not comparable with later Census figures, because the question was broadened in November 2025 from use in producing goods or services to use in any business function. Census working papers carry the standard disclaimer that they have not undergone the review accorded Census Bureau publications.
Peer-reviewed#
Eloundou, T., Manning, S., Mishkin, P. and Rock, D. (2024). GPTs are GPTs: Labor market impact potential of LLMs#
Science, 384(6702), 1306-1308
- Method
- Human and model ratings of task exposure across occupational task descriptions.
- Finding
- Around 80 percent of US workers could have at least 10 percent of tasks affected; about 19 percent could see at least half affected.
- What it supports
- Where pressure is likely to fall across the occupational structure.
- What it does not support
- That any job will be lost. This is exposure, not displacement, and the authors say so explicitly. It is the most misquoted number in the field.
Peer-reviewed#
Autor, D. and Thompson, N. (2025). Expertise#
NBER Working Paper 33941; Journal of the European Economic Association, 23(4), 1203-1271
- Method
- Four decades of task data across 303 US occupations, 1980-2018, with a novel content-agnostic measure of task expertise.
- Finding
- Automation that removed the LESS expert tasks raised wages and reduced employment. Automation that removed the EXPERT tasks lowered wages and increased employment.
- What it supports
- That which tasks are automated matters more than how many, and gives a testable way to ask whether a given role is appreciating or commoditising.
- What it does not support
- Anything measured about generative AI. The data ends in 2018, so this is a lens, not a forecast.
Peer-reviewed#
Babina, T., Fedyk, A., He, A. and Hodson, J. (2025). Firm Investments in Artificial Intelligence Technologies and Changes in Workforce Composition#
Chapter 3 in Technology, Productivity, and Economic Growth, NBER Studies in Income and Wealth 83, University of Chicago Press, July 2025. Earlier version NBER Working Paper 31325
- Method
- Worker resume and job-posting datasets combined to measure firm-level AI investment against workforce composition variables including educational attainment, specialisation and hierarchy.
- Finding
- AI investments are associated with a flattening of firms' hierarchical structure, with significant increases in the share of workers at the junior level and decreases in the shares in middle-management and senior roles.
- What it supports
- That the composition of the flattening, and not only its existence, is measurable, and that the layer shrinking is the one that has historically developed juniors into seniors.
- What it does not support
- Any magnitude. The abstract read for this entry states direction and significance and no percentage, so none should be attributed to it. Association rather than causation, and firms that invest in AI differ from those that do not in many other ways.
Working paper#
Ewens, M. and Giroud, X. (2025). Corporate Hierarchy#
NBER Working Paper 34162, issued August 2025, revised October 2025. DOI 10.3386/w34162. Not peer reviewed
- Method
- A measure of corporate hierarchy for over 3,100 US public firms, built from online resumes of 7 million US-based workers, 2016 to 2023, using a network estimation technique to identify hierarchical layers. AI adoption is proxied by AI job postings following Babina, Fedyk, He and Hodson.
- Finding
- Firms average ten hierarchical layers and a pyramidal structure, with the average and median number of layers declining across the sample period. More hierarchical firms show a more educated workforce, higher internal promotion rates, longer tenure, higher operating performance and higher administrative costs. Companies flattened their hierarchies following adoption of AI technologies, while pharmaceutical companies added layers after Covid-19. The authors state the AI tests are under-powered and that point estimates are significant at the 10 per cent level regardless of the adoption metric.
- What it supports
- That the Garicano prediction has now been tested against firm-level data and points in the predicted direction, which is more than the flattening discourse previously had.
- What it does not support
- That AI causes flattening. The authors call their own tests under-powered at the 10 per cent level, adoption is measured by job postings rather than by use, hierarchy is inferred from self-reported resumes, and the sample is US public firms. Not peer reviewed. Figures of 2,500 firms and 16 million employees circulate from an earlier draft and are wrong for this version.
Institutional survey#
Georgetown University Center on Education and the Workforce (2025). The Major Payoff: Evaluating Earnings and Employment Outcomes Across Bachelor's Degrees#
Georgetown CEW
- Method
- American Community Survey earnings and employment data across 152 majors for prime-age workers and 142 for early career.
- Finding
- Median prime-age earnings run from 58,000 dollars in education and public service to 98,000 in STEM. Within STEM alone the range is 64,000 to 146,000, and several humanities majors beat the STEM 25th percentile.
- What it supports
- That the spread within a field is often wider than the gap between fields, which undercuts advice given at the level of STEM against humanities.
- What it does not support
- Anything about the future. It is a cross-sectional snapshot of people already employed.
Working paper#
Humlum, A. and Vestergaard, E. (2025). Still Waters, Rapid Currents: Early Labor Market Transformation under Generative AI#
NBER Working Paper 33777, revised March 2026
- Method
- Adoption surveys linked to administrative labour records, roughly 25,000 workers across 7,000 Danish workplaces in 11 exposed occupations.
- Finding
- Precise null effects on earnings and hours two years after ChatGPT, ruling out effects larger than 2 percent, alongside substantial task reorganisation and new tasks in AI oversight and integration.
- What it supports
- That the structure of work moves well before earnings do, and that pay is the slowest available indicator.
- What it does not support
- That the same holds elsewhere. Denmark is high-trust, high-wage and heavily unionised, and two years is early.
Institutional survey#
Institute of Student Employers (2025). Student Recruitment Survey 2025#
Institute of Student Employers, 2025, with 2026 outlook published 7 January 2026
- Method
- Trade association survey of 155 employers covering over 31,000 student hires from more than 1.8 million applications in the 2024-2025 cycle.
- Finding
- An average of 89 applications per vacancy. The ISE reports a projected 7 per cent drop in student vacancies for 2026, while 30 per cent of employers increased student hiring.
- What it supports
- That the contraction is uneven. Nearly a third of employers increased student hiring in a falling market, which cuts against any account of uniform collapse.
- What it does not support
- Whole-market figures. It is a membership survey, and its members are organisations that run structured student recruitment in the first place.
Institutional survey#
Allen, J. S. (2026). Monitoring AI Adoption in the U.S. Economy#
FEDS Notes, Board of Governors of the Federal Reserve System, 3 April 2026. DOI 10.17016/2380-7172.4032
- Method
- Comparison of three independent US adoption measures: the Census Business Trends and Outlook Survey (firm-level, around 20,000 responses per wave), the Real-Time Population Survey (individual-level, 5,000 to 6,000 responses) and the Atlanta Fed Survey of Business Uncertainty (senior leaders, 1,032 responses). Four-period moving averages used for all BTOS calculations.
- Finding
- About 18 per cent of firms had adopted AI as of year-end 2025 on the BTOS. Work-related generative AI adoption in the RPS stood at about 41 per cent of the workforce as of November 2025, with daily use at 12 per cent. The SBU gives an employment-weighted firm adoption rate of about 78 per cent and an LLM adoption rate of about 54 per cent. Allen attributes the variation mainly to differences in sampling distributions and units of analysis, with question framing, the materiality of reported usage, information asymmetries and social desirability bias also contributing, and states senior leaders may face pressure to report AI usage as an efficiency initiative. The Census Bureau broadened its question in November 2025; the do-not-know rate ran at 10 to 11 per cent.
- What it supports
- That headline AI adoption figures differing by sixty points can all be correct, because they measure different units, and that the choice of measure decides the answer before any analysis begins.
- What it does not support
- Any productivity or employment effect. The note explicitly does not estimate AI's contribution to output, GDP or productivity, and names those as open questions beyond its scope. It also makes no claim that adoption has plateaued; the only slowdown language is deceleration in the second quarter of 2025.
Working paper#
Bonney, K., Breaux, C., Dinlersoz, E., Foster, L., Haltiwanger, J. and Pande, A. (2026). The Microstructure of AI Diffusion: Evidence from Firms, Business Functions, and Worker Tasks#
US Census Bureau, Center for Economic Studies Working Paper CES-26-25, April 2026
- Method
- Nationally representative data from the 2026 AI supplement to the US Census Bureau's Business Trends and Outlook Survey, analysed at three layers: overall firm use, deployment across business functions, and worker-task use. Reference period November 2025 to January 2026.
- Finding
- 18 per cent of firms used AI in a business function, rising to 32 per cent employment-weighted, with adoption expected to reach 22 per cent within six months. Use rates reach 50 to 60 per cent, and 60 to 70 per cent employment-weighted, for very large firms in Information, Professional Services and Finance. Among adopters, 57 per cent integrate AI in three or fewer business functions, most commonly Sales and Marketing (52 per cent), Strategy and Business Development (45 per cent) and IT (41 per cent). Workers use AI in work-related tasks in 23 per cent of firms, 41 per cent employment-weighted, and 65 per cent of firms limit use to three or fewer tasks. Most users, 66 per cent, rely on AI solely to augment tasks, and AI-related employment decreases occur in only 2 per cent of firms. Regression shows a positive correlation between firm commercial performance and the breadth of AI integration, holding across functional deployment, task-level use and operational investment; functional breadth and operational investment are positively associated with employment decreases, while worker-task integration shows no significant link to headcount reduction once the other two are accounted for.
- What it supports
- That AI diffusion is highly uneven by firm size and sector, and shallow even among adopters, which contradicts the premise that competitors hold equivalent capability.
- What it does not support
- Causation in either direction between adoption and performance: this is a cross-section of firms, and better-run firms may simply adopt more. Also not a stable time series. The Census Bureau broadened the underlying question in November 2025 from use in producing goods or services to use in any business function, which moved the level. A working paper, not peer reviewed, and US-only.
Working paper#
Brynjolfsson, E., Chandar, B. and Chen, R. (2026). Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence#
Stanford Digital Economy Lab, updated August 2026
- Method
- ADP payroll microdata covering millions of US workers, comparing employment by age and by occupational AI exposure since the release of ChatGPT.
- Finding
- No widespread economy-wide displacement. But employment among 22 to 25 year olds in highly AI-exposed occupations sits about 19 percent below where it would be had it tracked similarly aged workers in less-exposed occupations. The underlying levels matter and are easily lost: employment of that age group in the two most exposed quintiles fell about 11 percent between November 2022 and June 2026, while the same age group in the three least exposed quintiles grew about 10 percent. The 19 is the distance between those two, not a fall of 19. The divergence runs through reduced hiring rather than increased separations, and declines concentrate in occupations where AI substitutes for human tasks; where it complements, employment is flat or rising, especially for experienced workers.
- What it supports
- That the entry-level effect is real, measurable in payroll data rather than inferred, and specific to substitution rather than to AI exposure as such.
- What it does not support
- Economy-wide job destruction, which the authors explicitly rule out on current evidence. Nor does it establish causation: youth hiring is sensitive to interest rates, cohort size and hiring freezes, and the design is observational. Two further cautions, both from the paper itself. First, the phrase 'below trend' is wrong and is how this figure is usually repeated: the comparison is with same-aged workers in less-exposed occupations, not with a historical trend line and not with older workers. Second, the figures in circulation are not a worsening time series. 13 and 16 percent come from an earlier regression estimator, 15 and 19 from the descriptive kept-pace measure the authors now prefer, so quoting them in sequence as though the effect grew misrepresents a change of method as a change in the world. Note also that the ADP figures widely cited alongside this paper are the same payroll data reported differently, not independent corroboration.
Institutional survey#
Dixon, J.C. (2026). I surveyed workers to see if AI had caused job losses and was surprised by the findings#
The Conversation, 27 August 2026. YouGov survey commissioned by the author, College of the Holy Cross. Described by the author as a study in progress and not peer-reviewed
- Method
- 1,250 employed US workers surveyed online by YouGov between 30 July and 4 August 2026, 25 questions, weighted to the employed US population on age, gender, race and education. Opt-in panel. The author discloses an AI consulting business.
- Finding
- About 3 per cent said they had lost a job to AI since 2023, against roughly 6 per cent who said they held a job that did not exist before AI and about 9 per cent reporting an AI-related promotion. Around 95 per cent said no, with 3 to 4 per cent unsure.
- What it supports
- That self-attributed AI job loss is rare among people currently in work, and that reported AI-created gain in the same sample runs at twice the reported loss. Useful as a counterweight to displacement figures drawn from payroll data, which cannot ask anyone why.
- What it does not support
- Displacement in the workforce, and the reason is structural rather than a matter of sampling error. Every respondent was employed when surveyed, in Dixon's own words 'none were jobless, whether due to AI or another reason'. Anyone displaced by AI and still out of work is excluded from the numerator and the denominator alike, so the 3 per cent counts only those who lost a job and have since found another. It is a floor among survivors, not an estimate of displacement. It is also not a measure of fear: the question asked what happened, not what respondents expect.
Institutional survey#
Federal Reserve Bank of New York (2026). The Labor Market for Recent College Graduates#
New York Fed, data through 2026 Q2
- Method
- Current Population Survey and ACS tracking of unemployment and underemployment by major since 1990.
- Finding
- As at the second quarter of 2026, unemployment among recent graduates runs around 5.6 per cent and underemployment around 42 per cent, the highest since 2020.
- What it supports
- That the graduate labour market has tightened measurably, and that underemployment is the larger number.
- What it does not support
- Attribution to AI. The series is descriptive and the Fed states it is not a forecast.
Institutional survey#
High Fliers Research (2026). The Graduate Market in 2026#
High Fliers Research, January 2026
- Method
- Annual survey of graduate recruitment at the UK's 100 leading graduate employers. An employer survey of a selected group, not an official statistic.
- Finding
- Graduate recruitment fell 5.1 per cent in 2025, after a 14.6 per cent drop in 2024 and a 6.4 per cent decrease in 2023, with a further 0.5 per cent decrease forecast for 2026. Graduate recruitment at these employers has fallen 24.5 per cent since 2022, the lowest level since 2012. Employers received on average 23 per cent more applications in the first half of the 2025-2026 season, with applications roughly doubling since 2023.
- What it supports
- That the graduate entry route at the UK's largest recruiters has contracted by roughly a quarter in three years while competition for what remains has roughly doubled.
- What it does not support
- Attribution to AI, which the survey does not establish, and it covers 100 leading employers rather than the whole labour market.
Working paper#
Jadhav, R. and Danve, J. (2026). The AI Skills Shift: Mapping Skill Obsolescence, Emergence, and Transition Pathways in the LLM Era#
arXiv 2604.06906, 8 April 2026
- Method
- Benchmarking of four frontier models (LLaMA 3.3 70B, Mistral Large, Qwen 2.5 72B, Gemini 2.5 Flash) across 263 text-based tasks covering all 35 skills in the US Department of Labor O*NET taxonomy, 1,052 model calls. Cross-referenced against the Anthropic Economic Index. Measures models, not people. Preprint.
- Finding
- Introduces a Skill Automation Feasibility Index. Mathematics scores 73.2 and programming 71.8 for automation feasibility; active listening scores 42.2 and reading comprehension 45.5. All four models converge to similar skill profiles within a 3.6-point spread. Reports that 78.7 per cent of observed AI interactions are augmentation rather than automation, and a "capability-demand inversion" in which the skills most demanded in AI-exposed jobs are those the models perform least well at.
- What it supports
- That measured model capability and stated employer demand point in different directions, which is a useful counterweight to displacement forecasts built on exposure scores alone.
- What it does not support
- What happens to human capability. The authors state their index "measures LLM performance on text-based representations of skills, not full occupational execution". No humans were studied.
Vendor research#
Massenkoff, M. and McCrory, P. (2026). Labor market impacts of AI: A new measure and early evidence#
Anthropic, 5 March 2026. Corrected 8 March 2026. Read at source 5 September 2026
- Method
- A new occupational exposure measure, observed exposure, built from three inputs: O*NET task lists for roughly 800 US occupations, Eloundou and colleagues' theoretical LLM task-exposure scores, and Anthropic's own Claude usage data from the Anthropic Economic Index, weighting automated over augmentative use. Employment outcomes come from the Current Population Survey, in a difference-in-differences comparison of the top exposure quartile against the 30 per cent of workers with zero measured exposure.
- Finding
- No systematic increase in unemployment for highly exposed workers since late 2022; the pooled difference-in-differences estimate is small and indistinguishable from zero. On hiring, the monthly job-finding rate for 22 to 25 year olds entering the most exposed occupations fell by about 14 per cent in the post-ChatGPT period against 2022. The stable comparator runs at about 2 per cent per month in less exposed occupations, and entry into the most exposed jobs falls by roughly half a percentage point. The authors' own qualification, in their words, is that this is 'just barely statistically significant'. No such decrease appears for workers over 25.
- What it supports
- That a slowdown in youth hiring into exposed occupations shows up in a second dataset and a second method, alongside Brynjolfsson and colleagues on ADP payroll data. Also that the harder claim, a rise in unemployment among exposed workers, does not appear: the authors estimate they could detect a differential increase of about one percentage point and see nothing.
- What it does not support
- The 14 per cent is routinely restated as a fall against low-exposure peers. It is not. The paper's own words are 'compared to that in 2022 in the exposed occupations', so it is a change over time within one group, which is a weaker claim than the comparison usually reported. Nor does it establish cause: the authors list three benign readings, that unhired young workers may be staying in existing jobs, taking different ones, or returning to study, and note that survey-measured job transitions are prone to mismeasurement. TWO FURTHER CAUTIONS. The treatment variable is built from Anthropic's own product telemetry and cannot be reconstructed from outside the company. And on 8 March 2026 Anthropic corrected Figure 7, the job-finding figure this number is taken from, which had reversed the labels between the top-quartile and zero-exposure groups. Anyone citing a version of this chart captured between 5 and 8 March 2026 is citing the reversed one.
Working paper#
Rohde, W. (AiSuNe Foundation) (2026). Short-Term Gain, Long-Term Fragility: AI Labor Substitution and the Erosion of Sustainable Capability#
SSRN abstract 6577818, written 20 April 2026, revised 27 April 2026; also arXiv 2605.27399
- Method
- Sole-authored conceptual synthesis, 19 pages, no new empirical data. The author's own words: "This paper is a conceptual synthesis rather than a new empirical study" and "the evidentiary strategy is selective and scoped". Preprint, not peer reviewed, no journal reference. Dates verified at the SSRN record on 28 August 2026; the arXiv abstract page could not be fetched, so the arXiv submission date remains unconfirmed and is not asserted.
- Finding
- Develops a mechanism of capability masking followed by capability erosion: AI output creates a persuasive appearance that organisational capability has been replaced while dependence on skilled human labour remains, supporting hiring restraint and deferred structural reform while costs accumulate. Frames the result as a stack of deferred obligations: technical debt in artifacts and systems, capability debt in the human layer that maintains them, and institutional debt in the wider structures that reproduce skill and resilience.
- What it supports
- That the capability-erosion argument has been reached independently, from software engineering and political economy rather than from organisational research, and that the masking-before-erosion sequence is not unique to this estate's account.
- What it does not support
- Anything measured. It is a preprint that argues from other people's empirical work rather than presenting its own, and its societal-scale claims about fragility and concentration of power are inference rather than finding. It is also the reason no claim of first use is made for the term capability debt on this site. Checked at source on 28 August 2026: Rohde does not claim to have coined capability debt either. The paper contains no claiming language for it, and what he does claim is a mechanism, "to identify and formalize a mechanism of capability masking and capability erosion".
Working paper#
Shao, Y., Zope, H., Jiang, Y., Pei, J., Nguyen, D., Brynjolfsson, E. and Yang, D. (2026). Future of Work with AI Agents: Auditing Automation and Augmentation Potential across the U.S. Workforce#
arXiv:2506.06576, v3 revised 1 February 2026
- Method
- Audio-enhanced mini-interviews with 1,500 US domain workers across 104 occupations, covering 844 tasks drawn from O*NET, paired with capability assessments from AI experts. Introduces the Human Agency Scale, H1 to H5, and sorts tasks into four zones by desire against capability.
- Finding
- Worker preferences diverge sharply from technical capability. Tasks fall into an Automation Green Light Zone, an Automation Red Light Zone where capability exists and workers do not want it used, an R&D Opportunity Zone and a Low Priority Zone. Human Agency Scale profiles vary widely by occupation, and the authors report early signals of core competencies shifting from information-focused skills towards interpersonal ones.
- What it supports
- That the automate-or-not framing is too coarse, and that there is a measurable, occupation-specific preferred level of human involvement which does not track what the technology can do.
- What it does not support
- Nothing about what happens to capability when a task is automated. It measures what workers WANT and what experts think is POSSIBLE, which are both stated positions rather than outcomes. A preprint, not peer reviewed, and US-only.
Institutional modelling#
Waltmann, B. (2026). New Estimates of the Impact of Undergraduate Degrees on Lifetime Earnings#
Institute for Fiscal Studies, commissioned by the Department for Education
- Method
- Administrative linkage of school, university and tax records for the whole 2002 English GCSE cohort, tracked to age 37 with earnings simulated to 67.
- Finding
- Average net lifetime return to a degree is around 100,000 pounds, with very large variation by subject. Medicine and economics exceed 400,000 pounds on average. Creative arts, philosophy and languages show low or negative average returns. Around 20 per cent of women and 30 per cent of men are projected to see a negative net return.
- What it supports
- That subject choice carries a far larger financial spread than the decision to attend at all.
- What it does not support
- What any subject will return to someone choosing today. The authors explicitly decline to model structural change including AI.
The seven human capabilities
The evidence behind curiosity, empathy, adaptability and the rest, including where it is thinner than the claims usually made for it.
Peer-reviewed#
Swan, G. E. and Carmelli, D. (1996). Curiosity and mortality in aging adults: A 5-year follow-up of the Western Collaborative Group Study#
Psychology and Aging, 11(3)
- Method
- Prospective cohort. 1,118 men, mean age 70.6 at baseline, with an ancillary sample of 1,035 women.
- Finding
- Higher curiosity at baseline was associated with survival at five-year follow-up, and state curiosity remained significant after adjustment for other risk factors.
- What it supports
- That the association survives controlling for medical risk factors in one large cohort.
- What it does not support
- That curiosity extends life. Residual confounding, in particular underlying health driving both curiosity and survival, cannot be excluded.
Peer-reviewed#
Edmondson, A. (1999). Psychological Safety and Learning Behavior in Work Teams#
Administrative Science Quarterly, 44(2)
- Method
- Multi-method field study of 51 work teams in a single manufacturing company.
- Finding
- Team psychological safety predicted learning behaviour, which in turn mediated the relationship with team performance.
- What it supports
- That the route from safety to performance runs through learning behaviour rather than directly.
- What it does not support
- Causation, or generalisation beyond one manufacturing firm. It is a correlational field study.
Peer-reviewed#
Pulakos, E. D., Arad, S., Donovan, M. A. and Plamondon, K. E. (2000). Adaptability in the Workplace: Development of a Taxonomy of Adaptive Performance#
Journal of Applied Psychology, 85(4)
- Method
- Content analysis of over 1,000 critical incidents drawn from 21 jobs, followed by scale validation.
- Finding
- Eight dimensions of adaptive performance, including handling emergencies, managing stress, solving problems creatively and dealing with uncertain situations.
- What it supports
- That adaptability is decomposable into observable behaviours rather than being a single trait.
- What it does not support
- That employers reward these dimensions, or that they transfer across every occupation. The taxonomy was derived, not tested against outcomes.
Compiled review#
Simons, T. (2002). The High Cost of Lost Trust#
Harvard Business Review, September 2002
- Method
- Survey of more than 6,500 employees at 76 Holiday Inn hotels in the United States and Canada, matched to hotel financial records.
- Finding
- A one-eighth point improvement in managers' behavioural integrity rating was associated with a 2.5 per cent increase in profitability, roughly 250,000 dollars a year for an average hotel. No other measured aspect of manager behaviour had as large an effect on profits.
- What it supports
- That whether managers are seen to mean what they say tracks measurable financial outcomes.
- What it does not support
- Causation. It is a correlational study within one hotel chain, published in a practitioner magazine rather than a peer-reviewed journal.
Peer-reviewed#
Orlitzky, M., Schmidt, F. L. and Rynes, S. L. (2003). Corporate Social and Financial Performance: A Meta-Analysis#
Organization Studies, 24(3)
- Method
- Meta-analysis of 52 studies, 33,878 observations.
- Finding
- A positive association between corporate social performance and financial performance, which the authors describe as bidirectional.
- What it supports
- That principled conduct and financial results are not in general tension.
- What it does not support
- Causation in either direction, and the relationship is not uniform across how social performance is operationalised.
Peer-reviewed#
Maddux, W. W. and Galinsky, A. D. (2009). Cultural Borders and Mental Barriers: The Relationship Between Living Abroad and Creativity#
Journal of Personality and Social Psychology, 96(5)
- Method
- Five studies with MBA and undergraduate participants, combining correlational designs with causal priming experiments.
- Finding
- Time spent living abroad predicted success on creative-insight tasks and creative negotiation outcomes. Time spent travelling abroad did not.
- What it supports
- That adaptation to living in another culture, rather than exposure to it, is what relates to creativity.
- What it does not support
- A specific effect size for the general population. The samples are business and undergraduate students.
Peer-reviewed#
Harrison, S. H., Sluss, D. M. and Ashforth, B. E. (2011). Curiosity adapted the cat: The role of trait curiosity in newcomer adaptation#
Journal of Applied Psychology, 96(1)
- Method
- Longitudinal field study of 123 newcomers across 12 call-centre organisations.
- Finding
- Specific curiosity predicted information-seeking from colleagues, which in turn was associated with more creative handling of customer problems.
- What it supports
- That curiosity operates through a behavioural mechanism, asking, rather than as a disposition on its own.
- What it does not support
- That curiosity can be trained into people, or that the effect holds outside newcomer adaptation.
Peer-reviewed#
Konrath, S. H., O'Brien, E. H. and Hsing, C. (2011). Changes in Dispositional Empathy in American College Students Over Time: A Meta-Analysis#
Personality and Social Psychology Review, 15(2)
- Method
- Cross-temporal meta-analysis of 72 samples of American college students, total N 13,737, from 1979 to 2009.
- Finding
- Empathic Concern fell by 48 per cent and Perspective Taking by 34 per cent across the period, with most of the decline after 2000.
- What it supports
- That self-reported dispositional empathy declined measurably in this population over three decades.
- What it does not support
- A cause. The authors speculate about individualism and media but test no mechanism. It is also American college students only.
Peer-reviewed#
Sadri, G., Weber, T. J. and Gentry, W. A. (2011). Empathic emotion and leadership performance: An empirical analysis across 38 countries#
The Leadership Quarterly, 22(5)
- Method
- 360-degree ratings of 6,731 mid to upper-level managers across 38 countries. Subordinates rated empathy, superiors rated performance.
- Finding
- Managers rated as more empathic by subordinates received higher performance ratings from their own superiors, with the effect moderated by national power distance.
- What it supports
- That the association holds at scale and across many national contexts.
- What it does not support
- That empathy training improves performance. The design is cross-sectional and correlational.
Peer-reviewed#
von Stumm, S., Hell, B. and Chamorro-Premuzic, T. (2011). The Hungry Mind: Intellectual Curiosity Is the Third Pillar of Academic Performance#
Perspectives on Psychological Science, 6(6)
- Method
- Path-model synthesis of prior meta-analytic correlation matrices. Component samples range from 608 to 28,471.
- Finding
- Intellectual curiosity predicts academic performance independently of intelligence and effort, which the authors describe as a third pillar.
- What it supports
- That curiosity carries predictive weight that conscientiousness and ability do not account for.
- What it does not support
- Causation. It is a correlational synthesis. The widely quoted figure of roughly 50,000 students comes from the accompanying press release rather than from the paper itself.
Peer-reviewed#
Huang, J. L., Ryan, A. M., Zabel, K. L. and Palmer, A. (2014). Personality and Adaptive Performance at Work: A Meta-Analytic Investigation#
Journal of Applied Psychology, 99(1)
- Method
- Meta-analysis of 71 independent samples, total N 7,535.
- Finding
- Emotional stability and ambition predict adaptive performance, and the pattern differs from predictors of routine task performance.
- What it supports
- That individual differences carry predictive weight for adaptive performance specifically.
- What it does not support
- That adaptive performance is a distinct construct. This paper assumes the construct from earlier taxonomy work rather than establishing it.
Peer-reviewed#
Mehta, R., Zhu, R. and Meyers-Levy, J. (2014). When Does a Higher Construal Level Increase or Decrease Indulgence? Resolving the Myopia versus Hyperopia Puzzle#
Journal of Consumer Research, 41(2)
- Method
- Multi-study laboratory experiments.
- Finding
- Where the self is focal, a higher construal level increases indulgence rather than reducing it, reversing the effect the earlier literature predicted.
- What it supports
- That the benefit of stepping back is conditional, and the condition is identifiable.
- What it does not support
- That distant-future thinking is generally counterproductive. The effect is moderated, not reversed outright.
Institutional survey#
United States Environmental Protection Agency (2015). Notice of Violation, Volkswagen Group#
US EPA enforcement record
- Method
- Regulatory enforcement notice.
- Finding
- Affected 2.0-litre vehicles emitted nitrogen oxides at up to 40 times the standard in normal driving while appearing compliant in laboratory testing.
- What it supports
- That the defeat device produced a measured gap between test and road conditions of that magnitude.
- What it does not support
- That the same multiple applies across the range. The separate 3.0-litre notice cited up to nine times.
Institutional modelling#
Barton, D., Manyika, J., Koller, T., Palter, R., Godsall, J. and Zoffer, J. (2017). Measuring the Economic Impact of Short-Termism#
McKinsey Global Institute
- Method
- Corporate Horizon Index applied to 615 large and mid-cap United States public companies, 2001 to 2014.
- Finding
- Revenue of long-term firms grew cumulatively 47 per cent more than other firms, earnings 36 per cent more, and their share prices recovered faster after the financial crisis.
- What it supports
- That a measurable long-term orientation tracks with stronger cumulative growth in this sample.
- What it does not support
- Causation, or application outside large United States listed companies. The index is a constructed measure, not an observed policy.
Compiled review#
Gino, F. (2018). The Business Case for Curiosity#
Harvard Business Review, September to October 2018
- Method
- Survey of more than 3,000 employees across a range of firms, reported in a practitioner magazine.
- Finding
- Around 92 per cent said curious people bring new ideas to their teams, while about 24 per cent reported feeling curious in their jobs regularly.
- What it supports
- That the stated value of curiosity and the felt experience of it diverge sharply inside organisations.
- What it does not support
- Any causal link between curiosity and a business outcome. Self-report, not peer-reviewed, and the sample is not nationally representative.
Peer-reviewed#
Howick, J., Moscrop, A., Mebius, A. et al. (2018). Effects of empathic and positive communication in healthcare consultations: a systematic review and meta-analysis#
Journal of the Royal Society of Medicine, 111(7)
- Method
- Systematic review and meta-analysis of 28 randomised trials, 6,017 patients in total. Seven of the 28 tested empathic communication specifically; the remainder tested positive-expectation messaging.
- Finding
- The seven empathy-specific trials showed a small improvement in pain, anxiety and satisfaction, SMD -0.18, 95 per cent CI -0.32 to -0.03.
- What it supports
- That empathic communication has a measurable effect on patient-reported outcomes.
- What it does not support
- A large clinical benefit. The authors describe the effect as small, and only a quarter of the pooled trials tested empathy rather than positive framing.
Institutional survey#
Lorenzo, R., Voigt, N., Tsusaka, M., Krentz, M. and Abouzahr, K. (2018). How Diverse Leadership Teams Boost Innovation#
Boston Consulting Group
- Method
- Survey of more than 1,700 companies across eight countries.
- Finding
- Companies with above-average management diversity reported innovation revenue 19 percentage points higher than below-average companies, 45 per cent of total revenue against 26 per cent.
- What it supports
- That reported diversity and reported innovation revenue move together at scale.
- What it does not support
- Causation. Nor is this audited financial data. Note that 19 percentage points is not the same as 19 per cent higher, a distinction frequently lost in citation.
Peer-reviewed#
Stillman, P. E., Fujita, K., Sheldon, O. and Trope, Y. (2018). From 'Me' to 'We': The Role of Construal Level in Promoting Maximized Joint Outcomes#
Organizational Behavior and Human Decision Processes, 147
- Method
- Four laboratory and online experiments, pooled N approximately 691.
- Finding
- Prompting a higher level of construal led participants to choose options that maximised joint outcomes, including where doing so reduced their own payoff.
- What it supports
- That the level at which a problem is framed changes whether people optimise for themselves or for the whole.
- What it does not support
- Field behaviour. These are economic games with student and online samples.
Peer-reviewed#
Schlaegel, C., Richter, N. F. and Taras, V. (2021). Cultural intelligence and work-related outcomes: A meta-analytic examination of joint effects and incremental predictive validity#
Journal of World Business, 56(4)
- Method
- Meta-analysis of 70 studies providing 80 independent samples, total N 18,359.
- Finding
- Cultural intelligence is moderately associated with work-related outcomes, with a reliability-corrected average effect of about .39.
- What it supports
- That the association is consistent across a large body of studies.
- What it does not support
- Causation, and it does not establish any single dimension as the strongest predictor. The paper is about joint effects across all four dimensions.
Peer-reviewed#
De Freitas, J., Uguralp, A. K., Oguz-Uguralp, Z., Paul, L. A., Tenenbaum, J. and Ullman, T. D. (2023). Self-orienting in human and machine learning#
Nature Human Behaviour, 7
- Method
- Behavioural experiments with 124 human players across custom games, benchmarked against deep reinforcement learning agents.
- Finding
- Humans were near optimal at working out their own position and capabilities after conditions were altered. The reinforcement learning baselines were far from optimal at the same task.
- What it supports
- That rapid self-orientation after an unexpected change is currently a human advantage over the algorithms tested.
- What it does not support
- General workplace adaptability. These are simple custom games, the sample is modest, and the comparison is against specific algorithms rather than all AI approaches.
Compiled review#
Storey, M.-A. (2026). From Technical Debt to Cognitive and Intent Debt: Rethinking Software Health in the Age of AI#
ACM Queue, preprint at arXiv 2603.22106, March 2026
- Method
- Conceptual synthesis by a Canada Research Chair at the University of Victoria, drawing on Naur (1985) on programming as theory building, Cunningham (1993) on technical debt, and the author's own empirical work with Starr on developer confusion. Illustrated with a single teaching anecdote rather than a study. Proposes a framework; presents no new data.
- Finding
- Proposes a triple debt model: technical debt lives in code, cognitive debt lives in people as the erosion of shared understanding across a team, and intent debt lives in artefacts as the absence of captured rationale, goals and constraints. Argues generative AI may reduce technical debt while accelerating the other two, because code can now be produced faster than a team can build the understanding needed to change it safely.
- What it supports
- That the debt metaphor has been extended into a structured framework by a serious researcher, and that cognitive debt now carries a team-level meaning distinct from the individual-level one in Kosmyna et al. The distinction is the author's own and she states it explicitly.
- What it does not support
- Anything empirical about prevalence or magnitude. It is a framework paper with an anecdote, and it says so. Its reference list also mis-cites Kosmyna et al. as a 2024 CHI workshop paper, which does not appear on the MIT Media Lab's own publications list for that author; the citation is not relied on here.
The institutional record
What the most-quoted reports say, how they were made, and what they can carry.
Institutional survey#
United States Federal Aviation Regulations (1974). 14 CFR 121.441, Proficiency checks#
Electronic Code of Federal Regulations
- Method
- Binding regulation.
- Finding
- A pilot in command must pass a proficiency check every 12 calendar months, and within every 6 calendar months either a proficiency check or an approved simulator course.
- What it supports
- That one profession has made recurrent tested practice a legal condition of continuing to work.
- What it does not support
- That the intervals are calibrated to measured decay curves. The regulation sets a minimum, not an evidence-based optimum.
Institutional survey#
United States Federal Aviation Regulations (1981). 14 CFR 121.542, Flight crewmember duties, the sterile cockpit rule#
Electronic Code of Federal Regulations
- Method
- Binding regulation, Docket 20661, 46 FR 5502, 19 January 1981.
- Finding
- No crew member may perform any duty during a critical phase of flight other than those required for safe operation. Critical phases include taxi, take-off, landing and all operations below 10,000 feet except cruise.
- What it supports
- That protected attention can be written into law as a condition rather than left to individual discipline.
- What it does not support
- Anything about automation or skill decay directly. It is a rule about distraction.
Peer-reviewed#
Keil, M. (1995). Pulling the Plug: Software Project Management and the Problem of Project Escalation#
MIS Quarterly, 19(4), 421-447
- Method
- Longitudinal exploratory single-case study of one IT project inside a large computer manufacturer, pseudonymised as CompuSys. 111 interviews across eight job functions, 19 observed meetings, more than 350 collected documents.
- Finding
- CONFIG, an expert system built to help sales representatives produce error-free configurations before quoting, ran for over a decade and was terminated at the end of 1992 after, in the author's words, tens of millions of dollars. Successive business cases put net present value at 43.9 million dollars in 1982, 55.7 million in 1985 and at least 41.1 million in 1987. Keil concludes escalation is promoted by a combination of project, psychological, social and organisational factors rather than by any one of them.
- What it supports
- That the canonical study of a technology programme nobody could stop is a study of an artificial intelligence programme. The de-escalation literature is not being applied to AI by analogy; it began there and the field forgot.
- What it does not support
- Any prevalence. N is one, the organisation is pseudonymised, and there is no comparison group. An Academia.edu machine-generated summary of this paper asserts the project absorbed 250 million dollars; the article says tens of millions, and the 250 figure in it is 250 billion, being 1994 total US IT applications spending.
Peer-reviewed#
Keil, M., Mann, J. and Rai, A. (2000). Why Software Projects Escalate: An Empirical Analysis and Test of Four Theoretical Models#
MIS Quarterly, 24(4), 631-664
- Method
- Survey of information systems audit and control professionals, designed to gather data on projects that did not escalate as well as those that did, testing four theories: self-justification, prospect, agency and approach-avoidance.
- Finding
- The authors state that between 30 and 40 per cent of all IS projects exhibit some degree of escalation. The completion effect derived from approach-avoidance theory gave the best classification, correctly classifying over 70 per cent of both escalated and non-escalated projects.
- What it supports
- That escalation is common enough to be a base condition rather than an exception, and that the best-supported explanation is the pull of finishing rather than the psychology of self-justification alone.
- What it does not support
- Anything about AI, and nothing about whether escalated projects should have been stopped: it shows their outcomes were worse. Some degree of escalation is a soft threshold, and the respondents are auditors reporting retrospectively rather than a random sample of projects. The published sample size sits in the full text, which could not be opened during the 4 September 2026 build; the prevalence figure is quoted from the abstract and pages citing it say so.
Peer-reviewed#
Montealegre, R. and Keil, M. (2000). De-escalating Information Technology Projects: Lessons from the Denver International Airport#
MIS Quarterly, 24(3), 417-447
- Method
- Longitudinal qualitative case study of the automated baggage handling system at Denver International Airport, used to induce a process model of de-escalation.
- Finding
- De-escalation runs as a four-phase process: problem recognition, re-examination of prior course of action, search for alternative course of action, and implementing an exit strategy. The authors note that while escalation is well researched, there has been comparatively little research on the process of breaking the cycle.
- What it supports
- That climbing back down has a describable structure, and that the hard step is the second one, because re-examining the prior course of action requires the person who chose it to permit the question.
- What it does not support
- How often de-escalation is attempted or succeeds. N is one, the model is inductive rather than tested, and the case is a physical baggage system in the 1990s. The authors describe four phases, not stages.
Peer-reviewed#
Hughes, M. (2011). Do 70 Per Cent of All Organizational Change Initiatives Really Fail?#
Journal of Change Management, 11(4), 451-464. DOI 10.1080/14697017.2011.630506
- Method
- Critical review of five separate published instances of the 70 per cent organisational change failure rate, tracing each to its stated source.
- Finding
- In the author's words: 'whilst the existence of a popular narrative of 70 percent organizational change failure is acknowledged, there is no valid and reliable empirical evidence to support such a narrative.'
- What it supports
- That one of the most repeated statistics in management has no traceable empirical basis, and that this was established in a peer-reviewed journal fifteen years ago and ignored.
- What it does not support
- That change programmes usually succeed. The finding is about the absence of evidence for a specific number, not about the true rate, which remains unmeasured.
Institutional survey#
The Joint Commission (2013). Sentinel Event Alert 50: Medical device alarm safety in hospitals#
The Joint Commission, Issue 50, 8 April 2013
- Method
- Analysis of the Joint Commission's own Sentinel Event database, supplemented by FDA MAUDE reports and ECRI hazard rankings. Voluntary reporting; no sampling frame.
- Finding
- 98 alarm-related events between January 2009 and June 2012, of which 80 resulted in death, 13 in permanent loss of function and five in unexpected additional care or extended stay. 94 of the events occurred in hospitals. Contributing factors recorded as alarm signals inappropriately turned off (36), absent or inadequate alarm system (30), alarm signals not audible in all areas (25) and improper alarm settings (21). Reports that clinicians may turn the volume down, turn the alarm off, or set it outside safe limits in response to the volume of signals. Cites 566 alarm-related patient deaths in FDA MAUDE between January 2005 and June 2010. Led to National Patient Safety Goal NPSG.06.01.01, phased from 1 July 2014.
- What it supports
- That a regulator has documented, with named contributing factors, people disabling a safety control because it fired too often, and has had to legislate who holds the authority to change or silence it.
- What it does not support
- The size of the problem. The Commission's own footnote states that reporting is voluntary, represents only a small proportion of actual events, and that no conclusions should be drawn about relative frequency or trend. The widely quoted estimate that 85 to 99 per cent of alarm signals do not require clinical intervention is quoted BY the Commission from AAMI Horizons, Spring 2011, which is not a Joint Commission measurement and was not read for this entry.
Institutional modelling#
Arntz, M., Gregory, T. and Zierahn, U. (2016). The Risk of Automation for Jobs in OECD Countries: A Comparative Analysis#
OECD Social, Employment and Migration Working Papers No. 189
- Method
- Task-based re-estimation across 21 OECD countries using PIAAC survey data, accounting for the heterogeneity of tasks WITHIN occupations rather than treating whole occupations as automatable.
- Finding
- 9 per cent of jobs automatable on average across 21 countries, ranging from 6 per cent in Korea to 12 per cent in Austria. The authors state the occupation-based approach 'might lead to an overestimation of job automatibility, as occupations labelled as high-risk occupations often still contain a substantial share of tasks that are hard to automate'.
- What it supports
- That the headline automation figure is highly sensitive to whether you model occupations or tasks, and that the difference is roughly fivefold on the same question.
- What it does not support
- That 9 per cent is correct and 47 per cent wrong. Both are model outputs resting on assumptions, and neither has been scored against what happened.
Institutional survey#
World Economic Forum (2018). The Future of Jobs Report 2018#
World Economic Forum, Geneva, September 2018
- Method
- Employer survey via the WEF membership community.
- Finding
- Set out expected skill demand to 2022, with analytical thinking and innovation, active learning and creativity leading the list.
- What it supports
- What the field expected in 2018. Its predictions are now checkable, and that is the reason for including it.
- What it does not support
- A representative picture of employers. Respondents are drawn from a self-selected membership network.
Compiled review#
Davies, S. C., Atherton, F., Calderwood, C. and McBride, M. (2019). United Kingdom Chief Medical Officers' commentary on screen-based activities and children and young people's mental health and psychosocial wellbeing#
Department of Health and Social Care, Office of the Chief Medical Officer, 7 February 2019
- Method
- Official commentary by the four UK Chief Medical Officers on a commissioned systematic map of reviews. No new data.
- Finding
- States that scientific research is currently insufficiently conclusive to support UK CMO evidence-based guidelines on optimal amounts of screen use or online activities, and that the research does not present evidence of a causal relationship between screen-based activities and mental health problems, noting the possibility that young people who already have mental health problems spend more time on social media. It recommends a precautionary approach anyway, separates screen time from internet content and from persuasive design as three distinct issues, and advises families on the basis that screen time can displace health-promoting activities.
- What it supports
- That the UK's most senior clinical advisers concluded in 2019 that the dose measure could not support guidance, and named displacement rather than duration as the mechanism worth managing.
- What it does not support
- That screen use is harmless. The CMOs explicitly state that the absence of evident causal effect does not mean there is no effect. It is a commentary on a map of reviews rather than primary research, and it predates the AI question entirely.
Statutory investigation#
National Transportation Safety Board (2019). Collision Between Vehicle Controlled by Developmental Automated Driving System and Pedestrian, Tempe, Arizona, March 18, 2018#
NTSB Highway Accident Report NTSB/HAR-19/03, PB2019-101402, case HWY18MH010
- Method
- Statutory accident investigation of a single fatal collision, with access to vehicle system data, operator records and the developer's internal procedures.
- Finding
- The automated driving system detected the pedestrian 5.6 seconds before impact and tracked her to the crash without ever classifying her correctly or predicting her path. The developer had disengaged the Volvo XC90's factory forward collision warning and automatic emergency braking during automated operation. On detecting an emergency the system entered a one-second period of action suppression, withholding braking while it verified the hazard or the operator took control, and NTSB record that no alert was given to the operator when action suppression was initiated. The system recognised an imminent collision 1.2 seconds before impact. Probable cause was determined as the operator's failure to monitor the driving environment while visually distracted by a personal phone, with contributing factors including inadequate safety risk assessment procedures, ineffective oversight of vehicle operators and lack of adequate mechanisms for addressing operators' automation complacency.
- What it supports
- That a design can name a human as the primary countermeasure in an emergency and simultaneously withhold the alert that would let them act as one, and that the stated reason for doing so was concern about false alarms.
- What it does not support
- That the absence of an alert caused the crash. NTSB determined the probable cause to be the operator's inattention and found she would likely have had sufficient time to react had she been attentive. One vehicle, one developer, one jurisdiction, developmental software from 2018.
Institutional survey#
World Economic Forum (2020). The Future of Jobs Report 2020#
World Economic Forum, Geneva, October 2020
- Method
- Employer survey, conducted during the first year of the pandemic.
- Finding
- Named critical thinking and problem solving as leading skills, and forecast large-scale reskilling need.
- What it supports
- What the field expected in 2020, including a pandemic-shaped view of remote work.
- What it does not support
- A clean read on AI. The 2020 edition is dominated by COVID-era disruption.
Institutional survey#
McKinsey and Company (2021). Defining the skills citizens will need in the future world of work#
McKinsey Public and Social Sector Practice, June 2021
- Method
- Online psychometric survey of 18,000 people across 15 countries, fielded 2019; 56 elements in 13 skill groups.
- Finding
- Identifies distinct elements of talent, the DELTAs, associated with employment, income and job satisfaction.
- What it supports
- An unusually large individual-level dataset on skills and outcomes, which is rare in this literature.
- What it does not support
- Anything about AI. It was fielded in 2019, before the generative-AI period entirely.
Institutional survey#
International Organization for Standardization (2023). ISO/IEC 42001:2023, Artificial intelligence management system#
ISO, December 2023
- Method
- Certifiable management system standard developed by ISO/IEC JTC 1/SC 42.
- Finding
- Specifies requirements for an AI management system on a Plan-Do-Check-Act structure, against which an organisation can be certified by a third party.
- What it supports
- That an auditable management standard for AI now exists and can be certified.
- What it does not support
- Legal compliance. Certification against ISO 42001 is not the same as meeting the EU AI Act, and the two are routinely conflated.
Institutional modelling#
NFER (2023). The Skills Imperative 2035: An analysis of the demand for skills in the labour market in 2035 (Working Paper 3)#
National Foundation for Educational Research, with the University of Sheffield, funded by the Nuffield Foundation, May 2023
- Method
- 161 skills from the US O*NET database mapped to UK occupational codes, combined with UK employment projections.
- Finding
- Projects rising demand for a set of essential employment skills in the UK to 2035.
- What it supports
- A rare UK-specific, independently funded, methodologically documented projection.
- What it does not support
- Precision. It maps US skill data onto UK occupations, and a revised working paper corrects coding errors in the underlying labour force survey.
Institutional survey#
National Institute of Standards and Technology (2023). AI Risk Management Framework 1.0#
NIST, 26 January 2023
- Method
- Voluntary framework organised around four functions: Govern, Map, Measure and Manage.
- Finding
- Provides a common structure and vocabulary for AI risk management, with a companion Playbook and a 2024 generative AI profile.
- What it supports
- That a shared vocabulary exists that a board is likely to recognise.
- What it does not support
- Compliance with anything. It is voluntary and confers no legal status.
Institutional modelling#
National Institute of Standards and Technology (2023). AI Risk Management Framework Playbook, MANAGE 2.4#
NIST AI Resource Center, AI RMF 1.0 Playbook
- Method
- Voluntary framework and playbook developed through public consultation. Guidance rather than measurement.
- Finding
- Requires mechanisms and assigned responsibilities to supersede, disengage or deactivate AI systems showing performance inconsistent with intended use, and names five triggering conditions: end of system lifetime; risks exceeding tolerance thresholds; mitigation beyond the organisation's capacity; feasible mitigations failing regulatory, legal or normative standards; and impending risk detected in monitoring for which timely mitigation cannot be implemented. Decision thresholds for bypass or deactivation are treated as part of continual monitoring, and organisations are encouraged to provide contingency options including redundant or backup systems.
- What it supports
- That an authoritative framework treats deployment as reversible and expects the reversal mechanism, its thresholds and its fallback to exist before they are needed.
- What it does not support
- That any organisation does this, or that doing it works. The AI RMF is voluntary guidance, not a standard with conformity assessment, and it contains no evidence about outcomes.
Institutional modelling#
OECD (2023). OECD Skills Outlook 2023: Skills for a Resilient Green and Digital Transition#
OECD Publishing, Paris, November 2023
- Method
- Secondary analysis of OECD data including PISA and PIAAC, not a new survey.
- Finding
- Analyses the skills required for green and digital transitions across member economies.
- What it supports
- A cross-national, methodologically transparent baseline on skills, from data collected to a documented standard.
- What it does not support
- Anything AI-specific and current. The underlying data collection predates the generative-AI period.
Institutional survey#
World Economic Forum (2023). The Future of Jobs Report 2023#
World Economic Forum, Geneva, April 2023
- Method
- Employer survey on expectations to 2027.
- Finding
- Analytical thinking leads, with creative thinking second, and a growing emphasis on self-efficacy skills.
- What it supports
- The first post-ChatGPT edition, published five months after launch.
- What it does not support
- Considered judgement on generative AI. It was fielded too early for that.
Institutional survey#
AI Verify Foundation and IMDA, Singapore (2024). Model AI Governance Framework for Generative AI#
IMDA and AI Verify Foundation, 30 May 2024
- Method
- National governance framework for generative AI, nine dimensions. Both published PDF versions read in full and searched at source.
- Finding
- Contains zero occurrences of "human oversight", "human-in-the-loop", "over-reliance", "automation bias" or "deskill". Human oversight is not among its nine dimensions. The nearest it comes is a note that "core skills such as creativity, critical thinking and complex problem-solving are important to helping people harness AI effectively", and its only use of "competency" concerns third-party auditors rather than the human overseer.
- What it supports
- The value here is the confirmed absence, and the trajectory it establishes. In twenty months the same issuing body went from a framework with no oversight language at all to one built around it that also names deskilling. That shift is documented and quotable.
- What it does not support
- That Singapore was indifferent to oversight in 2024; the 2020 framework it builds on was not read at source and may carry such language. It establishes what the generative AI framework does not say, not what the whole regime did not say.
Peer-reviewed#
Chan, A., Ezell, C., Kaufmann, M., Wei, K., Hammond, L., Bradley, H., Bluemke, E., Rajkumar, N., Krueger, D., Kolt, N., Heim, L. and Anderljung, M. (2024). Visibility into AI Agents#
Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (FAccT 24), Rio de Janeiro, 3-6 June 2024. DOI 10.1145/3630106.3658948
- Method
- Conference paper assessing three categories of measure for increasing visibility into deployed AI agents, across a spectrum of centralised to decentralised deployment contexts and accounting for hardware and software service providers in the supply chain. Analysis and proposal rather than measurement.
- Finding
- Defines visibility as information about where, why, how and by whom AI agents are used, and assesses agent identifiers, real-time monitoring and activity logging as measures. Names five agent-specific risks: malicious use, overreliance and disempowerment, delayed and diffuse impacts, multi-agent risks, and sub-agents. On the last, the authors state that stopping an agent may require intervening on its sub-agents and that this may be difficult because, in their words, we lack methods for determining when an agent has created a sub-agent. The paper explicitly does not advocate immediate implementation of the measures and discusses their privacy and concentration-of-power costs.
- What it supports
- That the basic precondition of managing an agent, knowing what is running and what it has spawned, is an unsolved technical problem rather than a governance oversight.
- What it does not support
- That any of the proposed measures works, or is proportionate. The authors state they are describing options for further study rather than recommending deployment, and the paper contains no empirical evaluation.
Institutional survey#
DIFC Commissioner of Data Protection (2024). Regulation 10 on Personal Data Processed through Autonomous and Semi-Autonomous Systems#
Dubai International Financial Centre, DIFC-DP-GL-23 Rev.03, updated 27 August 2024; regulation enacted September 2023
- Method
- Binding regulation within the DIFC free zone, with accompanying guidance. Applies to personal data processing by autonomous and semi-autonomous systems, not to AI generally.
- Finding
- States that "human-defined processing purposes must always prevail in Systems development and use". Draws an explicit analogy between an autonomous system and an employee: where a system operates for the benefit of its deployer, "its position is substantially similar to that of an employee within the Deployer organization, and the Deployer should be therefore liable for its actions in the same way it may be liable for an employee's actions". Creates a named Autonomous Systems Officer performing a function similar to a data protection officer.
- What it supports
- That the employee analogy for accountability, and a named human role responsible for it, exist in a binding instrument somewhere in the Gulf.
- What it does not support
- Anything about capability or competence. It is a data protection regulation confined to one free zone, it addresses liability rather than skill, and it imposes no requirement that the responsible human be able to do the work being supervised.
Institutional modelling#
European Parliament and Council of the European Union (2024). Regulation (EU) 2024/1689, Annex III: High-risk AI systems referred to in Article 6(2)#
Official Journal of the European Union, official version of 13 June 2024. Text read via the European Commission AI Act Service Desk
- Method
- Binding legislative text. Annex III lists the eight areas in which AI systems are classified as high risk under Article 6(2), triggering the requirements of Chapter III.
- Finding
- Point 3 classifies as high risk AI systems used to determine access or admission to education, to evaluate learning outcomes including where those outcomes steer a learner's path, to assess the level of education a person will receive, and to monitor and detect prohibited behaviour during tests. Point 4 covers employment, workers' management and access to self-employment, naming systems used for recruitment or selection, including placing targeted job advertisements, analysing and filtering applications and evaluating candidates, and systems making decisions on promotion or termination, allocating tasks by individual traits, or monitoring and evaluating performance.
- What it supports
- That two of the applications discussed most loosely in public debate, marking pupils' work and sifting job applicants, are named in binding European law as high risk, with documentation, record-keeping, human oversight and explanation obligations attached.
- What it does not support
- Compliance or effect. Classification is not evidence that any system is biased or unsafe, and the Annex says nothing about how well the resulting obligations are met in practice.
Institutional modelling#
European Parliament and Council of the European Union (2024). Regulation (EU) 2024/1689, Article 13: Transparency and provision of information to deployers#
Official Journal of the European Union, official version of 13 June 2024. Text read via the European Commission AI Act Service Desk
- Method
- Primary legal text.
- Finding
- Requires high-risk systems to be designed so their operation is sufficiently transparent for deployers to interpret the output and use it appropriately, and to be accompanied by instructions for use stating the level of accuracy and its metrics, robustness and cybersecurity against which the system was tested and validated, plus known and foreseeable circumstances affecting that expected level, and where applicable information enabling deployers to interpret the output. Article 15(3) requires declared accuracy levels in those instructions.
- What it supports
- That European law requires documentary, aggregate, ex ante disclosure of accuracy and limitations to the deployer.
- What it does not support
- Any requirement to communicate uncertainty at the point of use. The word uncertainty appears in neither Article 13, Article 15 nor Article 50, and nothing requires a system to tell the person in front of it how confident it is in the specific output. EUR-Lex returned an empty document to every route tried on 4 September 2026, so the text was read at the Commission's own service desk; the Article 50 page there carries a notice that the provision has been amended by the Digital Omnibus and the displayed text not yet updated.
Institutional survey#
European Union (2024). Article 12, Record-keeping, Regulation (EU) 2024/1689#
Official Journal of the European Union
- Method
- Binding regulation.
- Finding
- High-risk AI systems must technically allow automatic recording of events over the system lifetime, to enable identification of risk situations, post-market monitoring and monitoring of operation. Article 19 requires providers to keep those logs for at least six months.
- What it supports
- That the capability to reconstruct what a system did is now a legal requirement rather than good practice.
- What it does not support
- That anyone must read the logs, or that a decision must be reviewable at the level of the individual case.
Institutional modelling#
European Union (2024). Regulation (EU) 2024/1689, Article 14: Human Oversight#
Official Journal of the European Union. Applies from 2 December 2027 for Annex III high-risk systems and 2 August 2028 for Annex I
- Method
- Binding regulation. Legal requirement rather than empirical finding.
- Finding
- Requires high-risk systems to be designed so they can be effectively overseen by natural persons, and requires that those persons be enabled to understand the system's capacities and limitations, to remain aware of the tendency to over-rely on its output (automation bias, named in the text), to interpret the output correctly, to decide not to use it or to disregard, override or reverse it, and to intervene or stop it through a stop button bringing the system to a safe halt. For biometric identification systems under Annex III point 1(a), no action may be taken unless the identification is separately verified by at least two natural persons with the necessary competence, training and authority.
- What it supports
- That an authoritative regulator treats override capability, not merely presence, as the content of oversight, and names competence, training and authority together in the one place it specifies who must verify.
- What it does not support
- That any of this happens. The oversight provisions do not apply until December 2027 at the earliest, the two-person requirement covers one narrow category rather than high-risk systems generally, and the regulation contains no evidence that oversight so constituted works.
Institutional modelling#
European Union (2024). Regulation (EU) 2024/1689, Article 26: Obligations of deployers of high-risk AI systems#
Official Journal of the European Union. Chapter III, Section 3
- Method
- Binding regulation. Legal requirement rather than empirical finding. Text read at source.
- Finding
- Paragraph 2 requires that deployers assign human oversight to natural persons who have the necessary competence, training and authority, as well as the necessary support. Paragraph 5 requires deployers to monitor operation on the basis of the instructions for use, to inform the provider and the relevant market surveillance authority without undue delay where the system presents a risk, and to suspend use of the system. Paragraph 6 requires retention of the automatically generated logs under the deployer's control for a period appropriate to the intended purpose and at least six months. Paragraph 7 requires employers, before putting a high-risk system into service at the workplace, to inform workers' representatives and the affected workers that they will be subject to its use. Paragraph 11 requires that natural persons subject to decisions made or assisted by an Annex III system be told.
- What it supports
- That the obligation to name a competent, authorised human overseer of a deployed system, and to be able to stop it, is law rather than good practice for systems in scope.
- What it does not support
- That it applies to most commercial agent deployments, which will fall outside the high-risk classification. Nor that any of it happens: the Regulation creates duties and does not evidence compliance. The application timetable has been subject to amendment and the published texts consulted for this entry did not agree on the dates, so no date is stated here.
Institutional modelling#
McKinsey Global Institute (2024). A new future of work: The race to deploy AI and raise skills in Europe and beyond#
McKinsey Global Institute, May 2024
- Method
- Modelling for 2022-2030 across nine EU countries, the UK and the US, plus a survey of 1,100 or more C-suite executives in five countries.
- Finding
- Projects large-scale occupational transitions and rising demand for social, emotional and higher cognitive skills.
- What it supports
- A transparent, well-documented scenario model, useful for direction.
- What it does not support
- What will happen. Scenario models are assumption-driven, and McKinsey's prior transition estimates have moved substantially between editions.
Institutional survey#
National Audit Office (2024). Use of artificial intelligence in government#
HC 612, Session 2023-24, 15 March 2024
- Method
- Value-for-money audit of the Cabinet Office and DSIT, including a survey of 87 government bodies conducted in autumn 2023, document review and interviews. Excludes simple rules-based automation, AI embedded by default in existing tools, and individuals' ad hoc use of public tools.
- Finding
- 37 per cent of responding bodies had deployed AI, typically one or two use cases; 70 per cent were piloting or planning, median four use cases. 21 per cent had an organisational AI strategy, with 61 per cent planning one. Of 32 bodies with deployed AI, 24 always or usually had a named accountable owner and 15 said use cases were always or usually identified at organisational level before deployment. 30 per cent of all respondents had risk and quality assurance processes explicitly incorporating AI risks. 70 per cent named difficulty recruiting or retaining AI skills as a barrier. The Cabinet Office's Central Digital and Data Office identified in 2023 that almost a third of civil service tasks, those it defined as routine, could be automated, and did not examine feasibility or assess cost.
- What it supports
- That the UK productivity claim for public sector AI rests on an indicative sizing exercise the auditor found untested for feasibility or cost, and that organisational ownership of deployed AI was incomplete in 2023.
- What it does not support
- The current position. The survey was taken in autumn 2023, before the generative wave reached most departments, and the picture will have moved. Survey response is self-reported and covers 87 bodies rather than the whole public sector.
Institutional survey#
Ryseff, J., De Bruhl, B. and Newberry, S. J. (2024). The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed#
RAND Corporation, RR-A2680-1
- Method
- Qualitative root-cause study. Interviews with 65 data scientists and engineers with at least five years building AI and machine-learning models in industry or academia.
- Finding
- A set of organisational anti-patterns behind AI project failure, chiefly misunderstood or miscommunicated intent, inadequate data, focus on the technology rather than the problem, missing infrastructure, and problems the technology cannot solve. The report opens by stating that by some estimates more than 80 per cent of AI projects fail, twice the rate of non-AI corporate IT projects.
- What it supports
- That the most repeated statistic about AI project failure does not come from a study. RAND footnote the 80 per cent to a press article and the comparison to a business magazine piece, and hedge it as by some estimates. Everyone quoting RAND drops the hedge.
- What it does not support
- Any base rate. This is 65 expert interviews about causes, not a measurement of frequency, and RAND do not claim otherwise. The two academic sources it cites on failure factors are themselves expert-interview studies rather than prevalence studies.
Peer-reviewed#
Wright, L., Muenster, R. M., Vecchione, B., Qu, T., Cai, P., Smith, A., COMM/INFO 2450 Student Investigators, Metcalf, J. and Matias, J. N. (2024). Null Compliance: NYC Local Law 144 and the Challenges of Algorithm Accountability#
Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, Rio de Janeiro, DOI 10.1145/3630106.3658998
- Method
- Field audit. 155 student investigators, acting as model job seekers, recorded the compliance of 391 employers with New York City Local Law 144 and the user experience for a prospective applicant. Accompanied by legal and policy analysis of the statute.
- Finding
- 18 employers posted a bias audit report, roughly 5 per cent, and 13 posted a transparency notice, roughly 3 per cent. The authors name the resulting state null compliance: non-compliance cannot be established because the law's design makes it impossible to determine whether an employer uses a covered tool. The analysis records that Local Law 144 requires an audit but is silent on its results, sets no discrimination threshold including the four-fifths convention, provides no remediation guidance, and that no federal safe harbour protects employers who disclose, so publication may create liability under other law.
- What it supports
- That the first algorithmic bias audit law in the world produced almost no public disclosure, and identifies a specific design reason for that rather than attributing it to employer indifference.
- What it does not support
- That the audited tools are biased or unbiased. No audit results were analysed because almost none were published, which is the finding. The sample is 391 employers with a large New York workforce rather than a census.
Institutional survey#
Court of Justice of the European Union (2025). Dun and Bradstreet Austria, Case C-203/22#
CJEU judgment of 27 February 2025
- Method
- Preliminary ruling interpreting GDPR Articles 15(1)(h) and 22.
- Finding
- A controller must describe the procedure and principles actually applied so the person can understand which of their data was used and how. Disclosing the algorithm is not a sufficient explanation, and a blanket trade-secret refusal is not permitted.
- What it supports
- That explanation now has a legal standard, and that the standard is comprehension rather than disclosure.
- What it does not support
- A right to source code or model weights. The balance with trade secrets is decided case by case.
Institutional survey#
Deloitte (2025). 2025 Global Human Capital Trends#
Deloitte Insights
- Method
- Around 10,000 business and HR leaders across 93 countries, plus separate worker, manager and executive surveys and 25 or more executive interviews.
- Finding
- Frames the worker-organisation relationship as a set of unresolved tensions rather than a set of solved problems.
- What it supports
- Scale and breadth of practitioner sentiment.
- What it does not support
- Causal claims. It is a sentiment survey by a firm that sells the remedies it recommends, and should be read with that in view.
Institutional survey#
Department for Education and IFF Research (2025). Technology in Schools survey: 2024 to 2025#
DfE research report, November 2025
- Method
- Survey of 1,634 schools in England, comprising 795 school leaders, 1,211 teachers and 489 IT leads, with qualitative interviews alongside. Questions on AI were new to the survey for 2025. Self-reported throughout.
- Finding
- 44 per cent of teachers reported using generative AI for school activities: lesson planning 35 per cent, delivering live lessons 7 per cent, marking 5 per cent. Teachers under 35 used it for planning at 43 per cent against 32 per cent for older colleagues, and for written feedback at 21 against 12 per cent. Teachers with under three years' experience used it for written feedback at 27 per cent against 14 per cent for those teaching longer. Leaders were more likely to plan investment in AI tools for teachers than for pupils, 58 against 20 per cent. Around one fifth of schools had a policy on safe and appropriate AI use. 77 per cent of secondary leaders whose pupils could access generative AI reported issues, most commonly plagiarism at 67 per cent.
- What it supports
- Where the teaching profession in England has actually placed the tool, on a large national sample: heavily in preparation, minimally in marking and delivery, and largely without written policy.
- What it does not support
- Any effect. It is a cross-sectional self-report of usage and perception, with no measurement of workload, learning or quality. The plagiarism figure records reported issues rather than the extent of plagiarism, which the report notes.
Institutional survey#
EY (2025). Work Reimagined Survey 2025#
EY Global, November 2025; UK cut published December 2025
- Method
- Employer and employee survey, 15,000 employees and 1,500 employers across 29 countries. Self-reported, cross-sectional, non-random sample. Published through a newsroom release rather than a methodology appendix.
- Finding
- 88 per cent of employees use AI at work, but mostly for basic tasks such as search and summarisation, and only 5 per cent use it in advanced ways that transform how they work. 37 per cent worry that overreliance on AI could erode their skills and expertise, rising to 43 per cent in the UK cut. Organisations pursuing AI gains on weak talent foundations saw productivity gains lag by over 40 per cent.
- What it supports
- That near-universal AI use coexists with very shallow use, and that employees themselves report skill-erosion worry at scale. It also gives a commercial argument for the capability case: the productivity is not collected where the human foundation is weak.
- What it does not support
- Any measured capability loss. Every figure is self-reported perception at a single point in time, the sample is not random, and the 40 per cent productivity figure is EY's own modelled comparison rather than an experimental result.
Institutional survey#
Internet Matters (2025). Me, myself and AI: Understanding and safeguarding children's use of AI chatbots#
Internet Matters, July 2025
- Method
- Mixed methods, March to July 2025. Survey of a representative sample of 1,000 UK children aged 9-17 and 2,000 parents of children aged 3-17, fielded April to May 2025; four focus groups with 27 children aged 13-17; 17 days of user testing across ChatGPT, Snapchat My AI and character.ai using two fictional child profiles; four expert interviews. Children are classified as vulnerable if they have an Education, Health and Care Plan, receive SEN support, or have a physical or mental health condition requiring professional help.
- Finding
- 64 per cent of children aged 9-17 have used an AI chatbot (ChatGPT 43 per cent, Google Gemini 32 per cent, Snapchat My AI 31 per cent). Among users the reasons are schoolwork 42 per cent, information 40 per cent, curiosity 40 per cent, chatting 24 per cent, advice 23 per cent, fun 18 per cent, wanting a friend 6 per cent, emotional help or therapy 3 per cent. 35 per cent say it feels like talking to a friend and 12 per cent say they use one because they have no one else to speak to, rising to 50 per cent and 23 per cent among vulnerable children, who are also nearly three times as likely to use companion-style products (17 against 6 per cent).
- What it supports
- That UK chatbot use among children is majority behaviour, dominated by schoolwork and information, and that companionship use concentrates in an identified vulnerable minority.
- What it does not support
- The precision of the vulnerable-group figures. Two charts are described as sharing the same base, children who have used at least one chatbot, but one uses 133 and 499 respondents and the other 188 and 802; the report does not flag this or publish significance testing, and its limitations section addresses the user testing only. The press release also substitutes a 42 per cent all-ages figure for the report's 47 per cent figure for 15-17 year olds. Nothing here is causal.
Compiled review#
Kolt, N. (2025). Governing AI Agents#
Notre Dame Law Review, Vol. 101, forthcoming. Preprint arXiv:2501.07913, submitted 14 January 2025, revised 11 February 2025
- Method
- Legal and economic analysis applying the economic theory of principal-agent problems and the common law doctrine of agency relationships to AI agents. No empirical component.
- Finding
- Characterises three problems arising from AI agents in agency terms: information asymmetry, discretionary authority and loyalty. Argues that the conventional solutions to agency problems, incentive design, monitoring and enforcement, might not be effective for governing AI agents that make uninterpretable decisions and operate at unprecedented speed and scale. Concludes that new technical and legal infrastructure is needed to support governance principles of inclusivity, visibility and liability.
- What it supports
- That a mature body of law and theory already exists for the question of who is accountable when something acts on your behalf, and that its standard remedies have identifiable failure points when the agent is artificial.
- What it does not support
- Anything about how agents behave in practice. It is a law review article arguing a position, cited here for its framework rather than as evidence of an outcome, and at the time of writing it is forthcoming rather than published.
Institutional survey#
Microsoft and LinkedIn (2025). 2025 Work Trend Index Annual Report: The Year the Frontier Firm is Born#
Microsoft WorkLab, April 2025
- Method
- 31,000 knowledge workers across 31 markets, plus LinkedIn labour data and Microsoft 365 telemetry.
- Finding
- Describes the emergence of firms organised around human-agent teams and a shift towards workers managing AI agents.
- What it supports
- Large-scale, current sentiment plus real product telemetry, which few others have.
- What it does not support
- Independence. Microsoft sells the tools whose adoption it is measuring, and telemetry measures usage rather than value.
Institutional modelling#
Ministry of Justice (2025). The use of evidence generated by software in criminal proceedings#
Call for evidence, 21 January to 15 April 2025, foreword by Sarah Sackman KC MP, Minister for Courts and Legal Services
- Method
- Government call for evidence, not a study. Sets out the current common law position, states the proposed boundaries of any reform, and puts five questions to respondents.
- Finding
- Records that section 69 of the Police and Criminal Evidence Act 1984, which required a party to show a computer was operating properly, was repealed on a 1997 Law Commission recommendation and replaced from 2000 by a common law rebuttable presumption that the computer was operating correctly at the material time. The foreword summarises this as the computer being always right unless someone shows otherwise, and cites the Post Office Horizon convictions as demonstrating the fallibility of software-generated evidence. Proposes that any reform cover evidence generated by software including artificial intelligence and algorithms, naming accounting systems, automated fraud and plagiarism detection and automated reporting from handheld devices, while excluding material merely captured by a device.
- What it supports
- That the legal presumption favouring machine output is live, is being reconsidered by the UK government, and that the government's own proposed scope for reform expressly includes AI and algorithmic systems.
- What it does not support
- Any outcome. It is a call for evidence rather than a decision, and no reform had been enacted at the date of review. It carries no data on how often the presumption is challenged or successfully rebutted.
Institutional survey#
Murray, A., House of Commons Library (2025). Apprenticeship statistics for England#
House of Commons Library briefing CBP 06113, 4 December 2025
- Method
- Parliamentary briefing compiling Department for Education apprenticeship statistics for England.
- Finding
- 353,500 apprenticeship starts in England in 2024/25, up from 340,000 in 2023/24 and 337,000 in 2022/23. In 2024/25, 51.3 per cent of starts were by apprentices aged 25 or over, 27.5 per cent aged 19 to 24, and 21.2 per cent under 19. The age distribution has remained approximately the same since 2018/19.
- What it supports
- That total apprenticeship starts are rising modestly and that half of all apprenticeships go to people aged 25 and over. The system is not primarily a route for the young and has not been for years.
- What it does not support
- Anything about quality, completion or whether apprenticeships lead to work. Starts are a count of beginnings.
Institutional modelling#
PwC (2025). PwC Global AI Jobs Barometer 2025#
PwC, 2025
- Method
- Analysis of close to a billion job advertisements across multiple countries.
- Finding
- Reports wage premiums for AI skills and shifting skill requirements in AI-exposed occupations.
- What it supports
- Large-scale observed labour demand rather than stated preference, which is a genuine strength.
- What it does not support
- Causation, and job advertisements describe what employers ask for rather than what the work requires.
Institutional survey#
Robb, M. B. and Mann, S. (Common Sense Media) (2025). Talk, Trust, and Trade-Offs: How and Why Teens Use AI Companions#
Common Sense Media, San Francisco, 16 July 2025. Fieldwork by NORC at the University of Chicago
- Method
- Survey of 1,060 US teens aged 13-17, interviewed 30 April to 14 May 2025, combining 719 probability interviews from NORC's AmeriSpeak Teen panel with 341 nonprobability interviews from Prodege, raked to February 2024 Current Population Survey totals. Margin of error plus or minus 4.2 percentage points. Cumulative response rate for the probability component 10.3 per cent.
- Finding
- 72 per cent have used an AI companion at least once and 52 per cent at least a few times a month. Against that: 80 per cent of users spend more time with real friends (68 per cent much more) and 6 per cent more time with AI; 67 per cent find AI conversations less satisfying than human ones; 50 per cent distrust the advice; 74 per cent have never shared personal information; 66 per cent have never felt uncomfortable; 9 per cent regard an AI as a friend or best friend; 46 per cent describe them as tools or programs.
- What it supports
- That conversational AI use is near-universal among US teenagers and that most of it is pragmatic rather than relational, on the report's own reading.
- What it does not support
- That 72 per cent use companion products. The definition given to respondents explicitly included using ChatGPT or Claude as companions, and the report's own limitations concede respondents may have conflated general AI use with companion use, potentially inflating usage statistics. It is cross-sectional and supports no causal claim, though the press release makes one. The published toplines and the report body also disagree on the social-skills transfer figure, giving 33 per cent and 39 per cent respectively.
Compiled review#
Stanford HAI (2025). The 2025 AI Index Report#
Stanford Institute for Human-Centered AI, April 2025
- Method
- Aggregation of many third-party sources across eight chapters, with public-opinion data from Ipsos and Pew.
- Finding
- The most comprehensive annual account of AI capability, investment, adoption and public attitudes.
- What it supports
- An authoritative baseline on what AI systems can do and how they are spreading.
- What it does not support
- Much about human capability. It measures the machine side of the equation. The human side is the gap this evidence base exists to fill.
Institutional survey#
State of Illinois (2025). Wellness and Oversight for Psychological Resources Act (HB1806), signed 1 August 2025#
Illinois Department of Financial and Professional Regulation
- Method
- State legislation, passed almost unanimously and signed by the Governor.
- Finding
- Prohibits the use of AI to provide therapy or perform therapeutic decision-making, including direct therapeutic communication with clients and detection of a client's emotional or mental state, while permitting administrative and supplementary use by licensed professionals. Penalties reach 10,000 dollars per violation.
- What it supports
- That at least one jurisdiction has moved from guidance to prohibition, and where it drew the line: the boundary is drawn at therapeutic decisions and direct therapeutic communication, not at the technology.
- What it does not support
- Anything about effectiveness, and nothing about other jurisdictions. One state, and its scope is contested.
Institutional modelling#
UNICEF Innocenti (2025). Guidance on AI and Children, Version 3.0: Recommendations for AI policies and systems that uphold child rights#
UNICEF Innocenti, December 2025
- Method
- Expert advisory group, multi-stakeholder consultation, peer review and a twelve-country study with children and caregivers.
- Finding
- Sets out ten requirements and 48 recommendations for AI systems and policy affecting children.
- What it supports
- A rights-based standard developed with children rather than about them, which almost nothing else in this space does.
- What it does not support
- Empirical claims about learning or capability effects. It is a normative guidance document.
Institutional survey#
United States Food and Drug Administration, Digital Health Advisory Committee (2025). Generative Artificial Intelligence-Enabled Digital Mental Health Medical Devices: meeting summary, 6 November 2025#
FDA Center for Devices and Radiological Health. Docket FDA-2025-N-2338
- Method
- Public advisory committee meeting with FDA presentations, 16 open public hearing speakers and structured committee deliberation on three scenarios: a prescription LLM therapy device for adults with major depressive disorder, over-the-counter and autonomous expansions, and use with under-21s.
- Finding
- The director of CDRH stated that FDA has authorised more than 1,200 AI-enabled medical devices and none yet involve generative AI for mental health conditions. The committee asked for premarket comparators BEYOND waitlist controls, judged autonomous over-the-counter use for undiagnosed users substantially higher risk, described multi-condition autonomous use as the highest risk of all, and expressed strong discomfort with autonomous use in children and adolescents. One member noted that reminders that a system is not human cannot overcome automation bias.
- What it supports
- The regulatory position as at November 2025, in the regulator's own words, and that the committee independently identified the waitlist-control weakness in the existing trial evidence.
- What it does not support
- What FDA will decide. Advisory committee recommendations are non-binding, and no rule follows from this meeting.
Institutional survey#
World Economic Forum (2025). The Future of Jobs Report 2025#
World Economic Forum, Geneva, January 2025
- Method
- Employer survey on expectations to 2030.
- Finding
- Analytical thinking is the most valued core skill, and skills gaps are named the single biggest barrier to business transformation.
- What it supports
- What employers say they want, which is a real and useful signal about demand.
- What it does not support
- What employers actually do. Stated skill preference and hiring behaviour diverge routinely.
Institutional modelling#
Acemoglu, D., Kong, D. and Ozdaglar, A. (2026). AI, Human Cognition and Knowledge Collapse#
NBER Working Paper 34910, February 2026, DOI 10.3386/w34910
- Method
- Dynamic theoretical model of learning and decision-making in which successful decisions require combining community-level general knowledge with individual context-specific knowledge, treated as complements. Human effort jointly produces a private signal and a thin public signal, creating a learning externality. Agentic AI substitutes for that effort. No empirical estimation.
- Finding
- Identifies a conditional tipping point: when human effort is sufficiently elastic and agentic recommendations exceed an accuracy threshold, the economy can reach a knowledge-collapse steady state in which general knowledge ultimately vanishes despite high-quality personalised advice. Welfare is non-monotone in agentic accuracy, implying an interior optimum. Greater capacity to aggregate and pool human-generated general knowledge raises welfare unambiguously.
- What it supports
- That the erosion argument can be stated formally with its assumptions visible, which is more than most of the vocabulary in this area manages. Also that the policy implication is not simply less AI: the model's unambiguous lever is better pooling of human knowledge, not lower agentic accuracy.
- What it does not support
- That knowledge collapse is happening or will happen. It is a model producing a possible steady state under stated conditions, it is a working paper rather than a peer-reviewed article, and the authors do not claim to have measured anything.
Institutional modelling#
Allianz Research (2026). Happy Labor Day? How geopolitics, immigration and AI will reshape work#
Allianz Trade, 30 April 2026
- Method
- Sectoral AI task-exposure estimates combined with national employment structures across the US, UK, Germany, France, Italy and Spain.
- Finding
- Models the combined effect of AI, demographics and migration on labour supply and task composition.
- What it supports
- A current, cross-country modelling view from outside the consultancy sector.
- What it does not support
- Measured effects. Like all exposure modelling, it is an estimate of what could be affected.
Institutional survey#
Boston Consulting Group (2026). When Everyone Uses AI, Companies Risk Losing Critical Skills#
BCG, 10 June 2026
- Method
- Global survey of 70 C-suite leaders and senior executives. Very small base, self-selected, and reporting perception rather than measurement.
- Finding
- Half of the executives surveyed report already observing deskilling in their organisations, and more than 60 per cent believe deskilling will pose a material threat to their organisation within the next three to five years.
- What it supports
- That deskilling has reached the point where senior leaders report seeing it themselves, which is a change in the executive agenda rather than in the evidence.
- What it does not support
- Prevalence. The base is 70 people. Anyone quoting the 50 per cent without the 70 is overstating it, and executive perception is not a measurement of what is happening to anyone's skills.
Compiled review#
Government Digital Service (2026). Algorithmic transparency records#
GOV.UK register maintained under the Algorithmic Transparency Recording Standard, first published January 2023, mandatory scope and exemptions policy December 2024. Count read 1 September 2026
- Method
- Public register of records completed by UK public sector organisations under the Algorithmic Transparency Recording Standard. Mandatory for all government departments and for arm's-length bodies delivering public or frontline services or interacting directly with the public; recommended for the wider public sector. Self-declared by publishing organisations.
- Finding
- 143 records published as at 1 September 2026, from central departments, agencies, local authorities, police forces and devolved administrations. Disclosed systems include a Department for Work and Pensions scanner reading around 25,000 scanned citizen documents a day to flag people who may need urgent assistance, a tool flagging Universal Credit journal messages that may indicate a risk of harm, the Cabinet Office verbal and numerical tests used to sift civil service applicants, an Ofsted tool drafting sections of children's home inspection reports, and adult social care case-note generation at a local authority.
- What it supports
- That algorithmic tools sitting between citizens and decisions about them are in production at volume in UK government, and that a public, structured record of some of them exists and can be read by anyone.
- What it does not support
- The extent of use. The register shows what has been disclosed and cannot show what has not. It is self-declared, the count moves, and the National Audit Office found in 2024 that the standard was not widely used before it became mandatory.
Institutional survey#
Infocomm Media Development Authority (IMDA), Singapore (2026). Model AI Governance Framework for Agentic AI, Version 1.0#
IMDA, published 22 January 2026, launched at Davos
- Method
- National governance framework for agentic AI, read in full at source. Builds on IMDA's 2020 Model AI Governance Framework. Guidance rather than statute. Law firms report a version 1.5 of 20 May 2026; that could not be confirmed at an official page, so version 1.0 is cited.
- Finding
- Names deskilling as a risk of agentic deployment, in the terms this research uses. Section 2.4.3: "As agents take over entry level tasks, which typically serve as the training ground for new staff, this could lead to loss of basic operational knowledge for the users. Organisations should identify core capabilities of each job and provide sufficient training and work exposure so that users retain foundational skills." Section 2.4 warns of "the potential loss of trade craft" and requires "sufficient training... to ensure that humans retain core skills". It also names automation bias directly, requires that overseers be trained to identify common failure modes, and requires that the effectiveness of human oversight itself be audited. It concedes that "continuous human oversight over all agent workflows becomes impractical at scale".
- What it supports
- That a national government has written the removal of entry-level work, and the consequent loss of the training ground for junior staff, into an operative AI governance framework. It is the closest external corroboration of the missing rungs argument found in any policy document.
- What it does not support
- Anything measured. It is guidance, not law, and it states a risk and a duty to train rather than evidence that deskilling has occurred. It also sets no threshold for what counts as retaining core skills, and no test of whether the training works.
Institutional modelling#
Leo XIV (2026). Magnifica Humanitas: Encyclical Letter on Safeguarding the Human Person in the Time of Artificial Intelligence#
The Holy See, given at Saint Peter's, 15 May 2026. 245 numbered paragraphs, 224 footnotes, five chapters
- Method
- Papal encyclical. Doctrinal and moral argument, read in full at source. Presents no original data and reports no study. It reasons from Catholic social teaching, explicitly continuing the line from Rerum Novarum (1891), whose 135th anniversary the signing date marks, through Laudato Si' and the 2025 Vatican note Antiqua et Nova.
- Finding
- States that AI use can weaken human capability, in terms specific enough to quote. Paragraph 100: heavy reliance and the search for ready-made answers can "weaken personal creativity and judgment". Paragraph 140: "every technology shapes those who use it", educating people about AI "involves teaching them to decide when and for what purpose it ought not to be used", and the ease of obtaining answers or summaries risks "extinguishing the desire to ask questions". Paragraph 150, quoting Antiqua et Nova, holds that "current approaches to technology can paradoxically de-skill workers, subject them to automated surveillance and relegate them to rigid and repetitive tasks". Paragraph 106 argues that "a slower pace in adopting AI does not mean opposing progress". Paragraph 156: "it is not enough to react only when jobs disappear; we must oversee the transformation in advance". Paragraph 198: "moral judgment cannot be reduced to calculation", and lethal or otherwise irreversible decisions may not be entrusted to artificial systems.
- What it supports
- That deskilling, the loss of the impulse to ask a question, and the case for deliberately slowing adoption are now stated in a magisterial document addressed to a global audience, rather than only in the research literature and the trade press. It is evidence about the standing of the argument, not about the world.
- What it does not support
- Nothing empirical whatsoever. It measures nothing, samples nobody and tests no hypothesis, and its deskilling claim is a quotation from an earlier Vatican note which is itself not an empirical study. Citing it as evidence that AI de-skills workers would be a category error. Its authority is moral and institutional.
Institutional modelling#
McKinsey Global Institute (2026). Agents, robots, and us: How AI reshapes work and skills in Europe#
McKinsey Global Institute, May 2026
- Method
- Task-level automation modelling across ten European economies covering more than 75 percent of regional labour force and GDP, plus job-postings data.
- Finding
- Models how agentic AI and robotics together reshape task composition and skill demand across Europe.
- What it supports
- The most current institutional modelling of the agentic shift, and useful for the direction of task change.
- What it does not support
- Outcomes. Task exposure modelling has consistently over-predicted the pace of realised change.
Institutional survey#
National Audit Office (2026). Increasing construction skills#
National Audit Office, 13 July 2026
- Method
- Audit of the government's construction skills package, including the foundation apprenticeships launched on 1 August 2025.
- Finding
- By April 2026 only 74 young people had started a construction foundation apprenticeship, against the department's own assumption of 1,000 in 2025-26.
- What it supports
- That the replacement scheme is not filling the gap it was designed for. A shortfall of this size, against the government's own planning assumption, in the sector the policy was built around, is the strongest available evidence that shortening apprenticeships has not by itself restored entry-level training.
- What it does not support
- That foundation apprenticeships cannot work. It is eight months of data on a scheme launched in August 2025, and low take-up in one sector.
Compiled review#
PwC (2026). Global AI Jobs Barometer 2026#
PwC, June 2026
- Method
- Analysis of close to a billion job advertisements across six continents, comparing skill requirements, wage premiums and job availability by occupational AI exposure.
- Finding
- Describes the labour market splitting into two paths, with human skills increasingly rewarded. The 2025 edition found a 56 percent wage premium for AI skills, jobs still growing in the most exposed occupations, and skills requirements changing 66 percent faster in the most exposed jobs.
- What it supports
- That demand-side signals in job advertisements are moving fast, and that the picture is not uniformly negative for exposed occupations.
- What it does not support
- Wage or employment outcomes for actual workers. Job advertisements are stated employer demand, not revealed price or realised hiring, and PwC has a commercial interest in the AI-skills market it is measuring.
Institutional modelling#
Skills England (2026). Evidence on defunding of level 7 apprenticeships#
Skills England, published 30 April 2026, produced March 2025
- Method
- Government evidence review supporting the decision to withdraw levy funding for level 7 apprenticeships for those aged 22 and over from 1 January 2026. Analysis of DfE apprenticeship statistics by level and age.
- Finding
- Level 7 apprenticeship starts in 2023/24 were 23,870, of which 65 per cent were aged 25 or over, 34 per cent aged 19 to 24, and 2 per cent under 19. Funding continues for 16 to 21 year olds, for care leavers and those with an education, health and care plan up to 24, and for anyone who started before 1 January 2026.
- What it supports
- That master's-level apprenticeship funding was overwhelmingly reaching people already established in careers rather than young entrants, which is the government's stated rationale. It is the clearest official statement of who higher-level apprenticeships actually served.
- What it does not support
- That degree apprenticeships in general were cut. Level 6, the undergraduate degree apprenticeship, is unaffected by this decision and its starts were still rising.
Professions and sectors
What AI does to expertise inside a named profession, and what the professional record shows.
Compiled review#
Transportation Safety Board of Canada (2019). Aviation Investigation Report A19P0112, Seair Seaplanes, Addenbroke Island#
Transportation Safety Board of Canada
- Method
- Single-accident investigation report. The passage on skill degradation sits in section 1.8 and reports concerns gathered from air-taxi operators in a separate TSB Safety Issue Investigation, not findings about this accident.
- Finding
- Surveyed air-taxi operators expressed concern that 'dependence on technology was causal in degradation of basic piloting skills', and numerous operators commented that over-reliance on GPS navigation 'may contribute to the decision to fly into adverse weather conditions'.
- What it supports
- That an industry regulator was recording practitioner-reported skill degradation from automation dependence, in an operational setting, before generative AI existed. It also captures the double bind: the sector's stated problem is too little technology and its observed failure mode is over-reliance on it.
- What it does not support
- That automation dependence caused this crash. The accident's own findings as to causes cite weather, terrain-alerting ambiguity and fatigue. This is industry survey commentary quoted inside an accident report, and citing it as a causal finding would misrepresent it.
Peer-reviewed#
Education Endowment Foundation and National Foundation for Educational Research (2024). ChatGPT in lesson preparation: a Teacher Choices trial#
EEF project record and evaluation report, 12 December 2024, project completed August 2026. Co-funded by the Hg Foundation; the ChatGPT guide was developed by Bain and Company's Social Impact practice
- Method
- Two-arm school-randomised Teacher Choices trial, 259 Year 7 and Year 8 science teachers across 68 state-funded secondary schools in England, ten weeks in the summer term of 2024. One arm used ChatGPT with a written guide; the other was asked to use no generative AI. Weeks one to five were a familiarisation period and planning time was recorded in weeks six to ten. Resource quality was assessed by an expert panel blinded to condition. Independent evaluation by NFER.
- Finding
- Weekly lesson and resource preparation time was 56.2 minutes in the ChatGPT arm against 81.5 minutes in the comparison arm, a saving of 25.3 minutes and a reduction of 31 per cent, given a high security rating. The blinded panel found no noticeable difference in resource quality. The proportion of ChatGPT-arm teachers who felt they spent too much time on preparation fell from 49 to 26 per cent, with no similar fall in the comparison arm. Frequency of use, and consultation of the guide, both declined over the trial.
- What it supports
- That light, largely unsupported use of a general assistant reduces teacher preparation time measurably, without a detectable cost to the quality of the materials produced, in one subject at one key stage.
- What it does not support
- Anything about teaching or about pupils. Preparation time was the outcome; classroom effect was outside scope. Time was self-recorded rather than observed. The sample over-represents schools in London and the South East and schools rated Outstanding, which the report states. The 31 per cent excludes a five-week familiarisation period.
Peer-reviewed#
Wilson, K. and Caliskan, A. (2024). Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval#
Proceedings of the 2024 AAAI/ACM Conference on AI, Ethics, and Society; preprint arXiv 2407.20371
- Method
- Resume audit study run through a document retrieval framework simulating job candidate selection. Massive Text Embedding models tested across nine occupations using over 500 publicly available resumes and over 500 job descriptions, with 120 first names associated with male, female, Black and white candidates. Code published by the authors.
- Finding
- The embedding models significantly favoured White-associated names in 85.1 per cent of cases and female-associated names in 11.1 per cent, with a minority of comparisons showing no statistically significant difference. Black male candidates were disadvantaged in up to 100 per cent of cases. Document length and the corpus frequency of a name also affected selection. Three hypotheses of intersectionality were validated.
- What it supports
- That the representation layer underneath commercial screening tools carries a large, measurable and intersectional name-based bias before any product logic is added, and that some of the effect is an artefact of name frequency in training data rather than anything about a candidate.
- What it does not support
- The behaviour of any deployed product. These are open embedding models in a simulated pipeline, not the proprietary systems vendors sell, which add filters and thresholds and cannot be independently tested. It measures ranking behaviour rather than who was hired.
Compiled review#
Divisional Court of England and Wales (Dame Victoria Sharp P and Johnson J) (2025). Ayinde v London Borough of Haringey and Al-Haroun v Qatar National Bank#
[2025] EWHC 1383 (Admin), judgment of 6 June 2025
- Method
- Two referrals under the Hamid jurisdiction, heard together, arising from fabricated citations placed before the court. In Ayinde, five cited authorities did not exist. In Al-Haroun, a claim for £89.4 million, eighteen of forty-five cited authorities did not exist.
- Finding
- At [6] the court held that freely available generative AI tools trained on a large language model are not capable of conducting reliable legal research. At [7] those using them have a professional duty to check accuracy against authoritative sources, which the court names. At [8] the duty extends to lawyers relying on others' AI-assisted work. At [81] a lawyer is not entitled to rely on their lay client for the accuracy of citations. At [23] the available powers run from public admonition and wasted costs to contempt and referral to the police, and at [31] admonishment alone is unlikely to suffice save in exceptional circumstances. At [9] leadership responsibility falls on heads of chambers and managing partners, and the court states it will inquire in future hearings whether that responsibility was fulfilled.
- What it supports
- That in England and Wales the verification duty is settled, non-delegable and extends upward to those who supervise. It also establishes that the court will treat a fabricated citation as a possible supervision failure rather than only as an individual one.
- What it does not support
- Anything about jurisdictions other than England and Wales, or about a lawyer who checked competently and was still misled, which no reported case has yet tested. It is a judgment rather than an empirical study, and it measures nothing.
Institutional modelling#
Financial Reporting Council (2025). AI in audit: Illustrative example and documentation guidance#
Financial Reporting Council, June 2025, published 26 June 2025
- Method
- First FRC guidance on artificial intelligence in statutory audit, developed with the FRC's Technology Working Group. Two parts: a worked example of an unsupervised machine learning tool used for journal risk assessment, and principles for documenting tools that use AI on the audit file.
- Finding
- The scope is deliberately broad, covering 'both traditional machine learning techniques and deep learning models, including generative AI'. On explainability the FRC declines to set a threshold: 'what constitutes appropriate explainability will vary widely based on context', and appropriate explanations 'may, particularly in relation to tools that rely on neural networks, be approximate or post hoc explanations that seek to explain how inputs influence outputs rather than the internal features and workings of the model'. Automation bias appears once, as something training material should carry 'strategies to mitigate'. Engagement teams are required to understand why the tool flagged an item and to stay alert to the possibility that its assessment is systemically flawed for that entity. The guidance states that it is not prescriptive and that 'the requirements against which firms will be assessed remain only those in the ISQMs and ISAs (UK)'.
- What it supports
- That a professional regulator can bring generative AI inside an existing evidence and documentation standard without writing new rules, and that it can do so while accepting post hoc explanation rather than model transparency.
- What it does not support
- Anything about practice. It is guidance issued in 2025, not a finding about what firms do, and the FRC says it creates no new requirements. It contains nothing on junior auditors, training pipelines or skills: the words junior, trainee and graduate do not appear.
Compiled review#
Financial Reporting Council (2025). Thematic Review: Certification of Automated Tools and Techniques#
Financial Reporting Council, June 2025, published 26 June 2025. Fieldwork Q2 2024 to autumn 2024
- Method
- Review of the processes and controls by which the six largest UK audit firms, named as BDO, Deloitte, EY, Forvis Mazars, KPMG and PwC, certify automated tools and techniques before use in audits. Information request issued April 2024, firm meetings summer 2024, feedback autumn 2024. A snapshot of process, not an inspection of audits.
- Finding
- All six firms had certification processes, but 'the maturity of these processes was found to vary and in some cases were not supported by formal documented policies'. Only two of the six set out the limitations of a tool or restrictions on its use in the certification documentation. Three captured assessment of the supporting IT control environment. One enforced a minimum recertification frequency, of three years. Generally the firms had no key performance indicators for tool usage and monitoring. The finding that carries furthest: 'There was no formal monitoring performed by the firms to quantify the audit quality impact of using ATTs.' At the time of review, generative AI use was limited to productivity aids such as chatbots rather than tools producing audit evidence.
- What it supports
- That the profession whose function is verification had, as at 2024, deployed the tools that produce its evidence without measuring their effect on the quality of that evidence, by its regulator's own account.
- What it does not support
- That audit quality has fallen. The review measures process rather than outcome, covers only the six largest firms, and its counts are of documentation practice rather than of tools or audits. Nothing here is a sample of engagements.
Compiled review#
Fletcher, J. and Verckist, D. (BBC and European Broadcasting Union) (2025). News Integrity in AI Assistants: An international PSM study#
BBC and European Broadcasting Union, published 21 October 2025
- Method
- Twenty-two public service media organisations across 18 countries and 14 languages put 30 shared news questions, drawn from questions audiences had actually asked, to the free consumer versions of ChatGPT, Copilot, Perplexity and Gemini between 24 May and 10 June 2025. Each prompt opened a new chat and asked the assistant to use the participating organisation's sources where possible. Assistants were anonymised and 271 journalists graded 2,709 core responses against accuracy, sourcing, separation of opinion from fact, editorialisation and context, marking each as no issues, some issues, significant issues or don't know.
- Finding
- Forty-five per cent of responses carried at least one significant issue, and 81 per cent carried an issue of some kind. Sourcing was the largest single cause at 31 per cent, then accuracy at 20 per cent and insufficient context at 14 per cent. Gemini recorded significant issues in 76 per cent of responses against 37 per cent for Copilot, 36 per cent for ChatGPT and 30 per cent for Perplexity, driven by sourcing, where Gemini's rate was 72 per cent against 24, 15 and 15. Of responses that cited a participating broadcaster's content, 15 per cent misrepresented it. Where the same BBC-only comparison could be run against the 2025 first round, significant issues fell from 51 per cent to 37 per cent.
- What it supports
- That misattribution and unsupported sourcing, rather than outright fabrication, is the dominant failure mode when general assistants answer news questions, and that it holds across languages, territories and platforms rather than being an artefact of one market or one model.
- What it does not support
- Current performance of any named product. These were free consumer versions tested in mid-2025 and all four have shipped new defaults since. It is not adversarial testing and difficulty was not controlled, so the rate is not a worst case; equally, per-organisation samples were around 120 responses, which the authors say is too small to compare countries or languages. Nothing here measures what readers then believed or did.
Institutional survey#
Government Digital Service (2025). Microsoft 365 Copilot Experiment: Cross-Government Findings Report#
Government Digital Service, June 2025
- Method
- Cross-government trial of M365 Copilot from 30 September to 31 December 2024, 20,000 licences across twelve organisations, each committing at least 1,000. Usage data from the Microsoft dashboard for 14,500 users; survey of 7,115 users; five focus groups. Time savings were self-estimated by selecting a band, and the average was computed from band midpoints with the largest savings estimated at 60 minutes.
- Finding
- Average self-reported saving 26 minutes a day. Adoption reached 83 per cent and held around 80 per cent. 17 per cent noticed no clear saving; more than a third reported over half an hour. Drafting documents 24 minutes, creating presentations 19, scheduling meetings 9. 82 per cent said they would not want to return to pre-Copilot conditions; satisfaction 7.7 and recommendation 8.2 out of 10; 85 per cent agreed it provided good value; 63 per cent believed their productivity would decline without it. The conclusions state it was not possible to identify how the saved time was spent.
- What it supports
- That a very large public sector deployment achieved high adoption and strongly positive user sentiment, and that accessibility benefits for disabled and neurodivergent users were a consistent theme.
- What it does not support
- Any measured productivity effect. The headline is a self-reported estimate, bucketed, with its top band capped by the analysts, and with no control group, no baseline task timing and no measure of output quality or decision quality. The report itself flags inconsistent user experience across departments and a festive-period disruption.
Peer-reviewed#
Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D. and Ho, D. E. (2025). Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools#
Journal of Empirical Legal Studies, 22(2), 216-242. DOI 10.1111/jels.12413
- Method
- First preregistered empirical evaluation of retrieval-augmented legal research tools. Over 200 handwritten legal queries across four categories, preregistered with the Open Science Foundation in March 2024, run against Lexis+ AI, Westlaw AI-Assisted Research, Ask Practical Law AI and GPT-4, graded on whether responses were correct and grounded in the sources cited.
- Finding
- The three commercial legal tools each hallucinated between 17 and 33 per cent of the time. Lexis+ AI was accurate on 65 per cent of queries and incomplete on 18 per cent; Westlaw AI-Assisted Research was accurate 42 per cent of the time with a hallucination in one-third of responses; Ask Practical Law AI was incomplete on 62 per cent of queries. Westlaw produced the longest answers, averaging 350 words against 219 and 175, which the authors link to both its higher hallucination rate and the verification burden it imposes.
- What it supports
- That retrieval-augmented generation reduces hallucination relative to a general chatbot without eliminating it, and that provider claims of hallucination-free citation were overstated. Longer generated answers carry more falsifiable propositions and require more checking.
- What it does not support
- Current performance of any named product. These are specific versions tested in 2024 and providers update continuously. It also does not measure legal outcomes, only response accuracy against expert grading.
Institutional survey#
Arzilli, F., Lynch, C. S. and Page, L. (2026). An Evaluation of DWP's Microsoft 365 Copilot Trial#
Department for Work and Pensions, published 29 January 2026
- Method
- Mixed-methods post-implementation evaluation of a trial running October 2024 to March 2025 with 3,549 licensed staff. Survey of users (1,716 responses) and of a random stratified comparison group of non-users (2,535 responses from 9,300 sampled), 19 qualitative interviews, and seemingly unrelated regression controlling for demographic, occupational and AI-keenness variables. No baseline; licences allocated first come, first served.
- Finding
- Estimated saving of 19 minutes a day across eight routine tasks, statistically significant across all specifications, with the largest task effects on searching for information (26 minutes), writing emails (25) and summarising (24). 90 per cent of users said it saved time. Job satisfaction rose 0.56 points and perceived work quality 0.49 points on seven-point scales; 73 per cent reported better quality outputs and 65 per cent felt more fulfilled.
- What it supports
- That a regression-based estimate on a large departmental sample, controlling for observable differences, still finds a positive and significant self-reported effect on efficiency, satisfaction and perceived quality.
- What it does not support
- A measured time saving. The evaluation's own limitations chapter names the absence of baseline data, post-treatment bias, self-selection towards AI enthusiasts through first-come first-served allocation which it says may lead to overestimation, non-response bias and acquiescence bias on the time question. Nothing about decision quality or citizen outcomes was measured.
Institutional modelling#
Board of Governors of the Federal Reserve System, Federal Deposit Insurance Corporation and Office of the Comptroller of the Currency (2026). Supervisory Guidance on Model Risk Management (SR 26-2)#
Federal Reserve supervisory letter SR 26-2, 17 April 2026, 12 pages. Supersedes SR 11-7 (4 April 2011) and SR 21-8 (9 April 2021)
- Method
- Interagency supervisory guidance for United States banking organisations, most relevant to those above $30 billion in total assets. Replaces the 2011 model risk management framework after fifteen years of supervisory experience.
- Finding
- The definition of a model is narrowed to 'a complex quantitative method, system, or approach that applies statistical, economic, or financial theories to process input data into quantitative estimates', expressly excluding simple spreadsheet arithmetic, deterministic rule-based processes and software with no such theory underpinning it. Footnote 3 states that generative and agentic AI models 'are novel and rapidly evolving' and 'are not within the scope of this guidance', while the principles do apply to traditional quantitative models and to non-generative, non-agentic AI. Effective challenge is retained and defined as critical analysis by objective experts with expertise, independence and the organisational standing to force change. The guidance sets no enforceable standards and says non-compliance will not itself draw supervisory criticism.
- What it supports
- That the most mature oversight regime any profession has for machine-produced numbers has, on its own initiative and in 2026, placed generative AI outside its scope. Anyone citing model risk management as the ready-made precedent for governing generative AI in finance is citing a document that declines the job.
- What it does not support
- That generative AI in banks is ungoverned. The same footnote directs firms to their own risk management and governance practices for tools outside scope, and other supervisory expectations, consumer protection law and third-party risk guidance still apply. It is United States banking supervision only, and guidance rather than rule.
Compiled review#
Charlotin, D. (2026). AI Hallucination Cases database#
damiencharlotin.com, updated daily. Figures read from the update of 27 August 2026
- Method
- Curated database of legal decisions worldwide in which a court or tribunal has explicitly found or implied that a party relied on hallucinated material. Excludes mere allegations, with a small stated exception. Coverage begins in the second quarter of 2023.
- Finding
- 1,963 cases identified as at 27 August 2026. By jurisdiction: United States 1,345, Canada 214, Australia 98, United Kingdom 62, Israel 57, with more than thirty other countries represented. By party responsible: self-represented litigants 1,127, lawyers 784, judges 29, expert witnesses 15. By nature: fabricated material 1,634, misrepresented authority 816, false quotations 528, outdated advice 33.
- What it supports
- That fabricated legal authority reaching courts is a documented, dated, worldwide phenomenon rather than an anecdote, and that self-represented litigants account for more recorded instances than lawyers do.
- What it does not support
- The true rate. The database counts decisions where a court addressed the point, so instances nobody noticed, or resolved without a written decision, are invisible by construction; its author states this. It is a curated compilation rather than a sampled study, so it cannot support a denominator or a trend rate.
Compiled review#
Kissin, E. (Financial Times) (2026). Junior consultants called back to office as AI increases need for human skills#
Financial Times, 27 August 2026
- Method
- News reporting. On-the-record interviews with named executives at EY, KPMG, Accenture and Azets, with reference to policies at Deloitte, PwC, BCG, Microsoft and JPMorgan. Not a study, and no measurement of skill or outcome.
- Finding
- Consulting leaders report that AI has raised the value of interpersonal skills and are considering requiring junior staff in the office more often to develop them. EY's UK head of consulting, Sayeh Ghanbari, is quoted saying firms will "have to reduce flexibility, but in order to help the human skills", and that firms dropped training in empathy, storytelling and leadership during the remote-working period while prioritising AI and technical skills. KPMG describes reinventing in-person training; BCG is expanding office social activities; Azets has lowered its degree requirement and is encouraging four days a week. Deloitte and PwC began extra coaching for their youngest UK recruits in 2023 after finding weaker teamwork and communication than earlier cohorts.
- What it supports
- That the apprenticeship-erosion argument is now being acted on by the largest professional services firms, named and on the record. It is strong evidence of institutional belief and of policy change.
- What it does not support
- That AI caused the deficit, or that office attendance repairs it. No measurement appears anywhere in the reporting, the cohort effects described in 2023 are attributed to pandemic lockdowns rather than to AI, and EY as a firm restated its existing flexibility policy alongside its executive's comments. Testimony from interested parties is not evidence of a mechanism.
Compiled review#
National Transportation Safety Board (2026). Highway Investigation Report HIR-26-02, Ford BlueCruise collisions#
National Transportation Safety Board, 31 March 2026
- Method
- Investigation of two fatal crashes in which Ford BlueCruise-equipped vehicles struck stationary vehicles at highway speed: San Antonio, 24 February 2024, and Philadelphia, 3 March 2024.
- Finding
- Overreliance on the partial automation system appears in the probable cause for both crashes, not merely in discussion: distraction 'stemming from overreliance on the vehicle's hands-free partial automation system' in San Antonio, and 'overreliance on and misuse of' it in Philadelphia. The board recommended that driver monitoring systems detect and warn about 'accumulated short distractions' over a prolonged period.
- What it supports
- That an investigator placed automation over-reliance in the causal chain rather than treating it as context, and that the monitoring system was defeated by ordinary attention behaving ordinarily rather than by anyone circumventing it.
- What it does not support
- Anything about knowledge work. Driving is a continuous manual-control task with a monitoring system watching the human, which is not the shape of AI-assisted professional judgement.
Frontline, clinical and physical work
What AI does to skill outside the office, where most of the world's work happens.
Peer-reviewed#
Cook, R. I., Render, M. and Woods, D. D. (2000). Gaps in the continuity of care and progress on patient safety#
BMJ, 320(7237), 791-794
- Method
- Conceptual paper in the education and debate section, developed against two publicly documented US clinical accidents.
- Finding
- Gaps are defined as discontinuities in care, appearing as losses of information or momentum or interruptions in delivery. Most gaps are anticipated and bridged by practitioners, so invisibly that neither outsiders nor insiders recognise the activity as distinct work. Bridging is not elimination: some bridges are frail and easily undone. Accidents occur when conditions overwhelm the mechanisms practitioners use to detect and bridge gaps, from which the authors conclude that efforts to forestall errors by isolating practitioners from the system will misfire. Their worked example is the division of nursing work with less credentialed technicians, which delivers a real economic benefit and restricts the nurse's ability to anticipate gaps.
- What it supports
- That the boundary between steps in a process is a named object with its own failure modes, and that the work of bridging it is invisible to every measure of output. The nursing example transfers directly to delegating work to agents.
- What it does not support
- Anything quantified. There is no sample, no rate and no measurement, and nothing in it concerns automation or AI. It is a framing paper.
Peer-reviewed#
Drew, B. J., Harris, P., Zegre-Hemsey, J. K., Mammone, T., Schindler, D., Salas-Boni, R. et al. (2014). Insights into the Problem of Alarm Fatigue with Physiologic Monitor Devices: A Comprehensive Observational Study of Consecutive Intensive Care Unit Patients#
PLoS ONE, 9(10), e110274. DOI 10.1371/journal.pone.0110274, published 22 October 2014
- Method
- Observational study of consecutive adults in five adult intensive care units at the University of California, San Francisco, over the 31 days of March 2013. All monitor data, seven ECG leads, pressure, SpO2 and respiration waveforms, user settings and alarms, were stored for 461 patients. Nurse scientists annotated 12,671 arrhythmia alarms against a defined protocol, with inter-rater agreement of 95 per cent on true against false and a Cohen kappa of 0.86. Funded by GE Healthcare.
- Finding
- 2,558,760 unique alarms occurred in 31 days: 1,154,201 arrhythmia, 612,927 parameter and 791,632 technical. 381,560 were audible, an audible alarm burden of 187 per bed per day. 88.8 per cent of the 12,671 annotated arrhythmia alarms were false positives, and 93 per cent of the 168 true ventricular tachycardia alarms were not sustained long enough to warrant treatment.
- What it supports
- That a safety control firing at this rate trains the person holding it to ignore it, and that the training is rational rather than negligent. It is the best measured case anywhere of the base-rate problem that makes a stop control nominal.
- What it does not support
- Anything about AI. The 88.8 per cent applies only to the 12,671 annotated arrhythmia alarms, NOT to all 2.56 million alarms and not to clinical alarms in general; a page stating that 88.8 per cent of clinical alarms are false has misread it. Single centre, one month, five units, industry funded.
Statutory investigation#
National Transportation Safety Board (2014). Descent Below Visual Glidepath and Impact With Seawall, Asiana Airlines Flight 214, Boeing 777-200ER, HL7742, San Francisco, California, July 6, 2013#
NTSB/AAR-14/01, adopted 24 June 2014
- Method
- Statutory accident investigation with access to flight recorders, crew interviews, manufacturer documentation and operator training records. N is one.
- Finding
- The pilot flying selected a mode that caused the autoflight system to climb, then disconnected the autopilot and moved the thrust levers to idle, which put the autothrottle into HOLD, a mode in which it does not control airspeed. The board records that neither the pilot flying, the pilot monitoring, nor the observer noted the change in autothrottle mode to HOLD. Contributing factors in the probable cause include the complexities of the autothrottle and autopilot flight director systems being inadequately described in the manufacturer's documentation and the operator's training, which increased the likelihood of mode error. The board attributes insufficient airspeed monitoring in part to automation reliance, and notes the operator's automation policy emphasised full use of automation and did not encourage manual flight in line operations.
- What it supports
- That a correct automation state change can constitute a failed handoff. Nothing malfunctioned: authority moved from machine to crew and the crew were not told in a way that reached them.
- What it does not support
- Any frequency of mode confusion. One accident, three crew, one aircraft type. It cannot support a rate and is used here as a demonstration of a mechanism.
Peer-reviewed#
Starmer, A. J., Spector, N. D., Srivastava, R., West, D. C., Rosenbluth, G., Allen, A. D. et al. for the I-PASS Study Group (2014). Changes in Medical Errors after Implementation of a Handoff Program#
New England Journal of Medicine, 371(19), 1803-1812
- Method
- Prospective systems-based intervention study across nine paediatric residency programmes in the US and Canada, January 2011 to May 2013, in three staggered waves with season-matched six-month pre and post periods. 875 consenting residents, 10,740 patient admissions. Active surveillance five days a week, incidents classified by two blinded physician reviewers. Not randomised, no control group.
- Finding
- The medical-error rate fell 23 per cent, from 24.5 to 18.8 per 100 admissions, and preventable adverse events fell 30 per cent, from 4.7 to 3.3 per 100 admissions, both P<0.001. Near misses and non-harmful errors fell 21 per cent. Non-preventable adverse events did not change, 3.0 against 2.8, P=0.79. Oral handoff duration did not change, 2.4 against 2.5 minutes per patient, P=0.55. Error rates did not change significantly at three of the nine sites, although written and oral handoff processes improved at all nine.
- What it supports
- That a handoff can be made substantially safer without being made longer, by changing what is transferred rather than how much time is spent transferring it. The unchanged non-preventable rate is the strongest internal evidence the effect is real.
- What it does not support
- Causation, by the authors' own statement, and not which element of the bundle did the work. Paediatric inpatient units only; generalisation to other specialties is untested. Reviewer agreement was moderate, kappa 0.47 for error classification. Data collectors could not be blinded to period. Funding included an unrestricted medical education grant from Pfizer.
Compiled review#
The Joint Commission (2017). Sentinel Event Alert 58: Inadequate hand-off communication#
The Joint Commission, Issue 58, 12 September 2017
- Method
- Advisory bulletin aggregating third-party findings. No original data, no sample, no denominator.
- Finding
- States that inadequate hand-off communication contributes to adverse events including wrong-site surgery, delay in treatment, falls and medication errors. Reports, from cited third parties, that communication failures were responsible at least in part for 30 per cent of US malpractice claims over five years, 1,744 deaths and 1.7 billion dollars in costs; that a typical teaching hospital may experience more than 4,000 hand-offs a day; and that 69 per cent of clinical learning environments had no standardised hand-off process against 20 per cent with some standardisation. Its only quantified outcome evidence is the I-PASS trial.
- What it supports
- That handoff has been named as a systemic failure point by a national accreditation body, and that the profession's own quantified evidence for it is thinner than the advisory framing implies.
- What it does not support
- The most quoted claim attached to it. The statement that 80 per cent of serious medical errors involve miscommunication during handoff does not appear anywhere in this alert; the full six pages were read on 4 September 2026. The nearest real figure is Starmer and colleagues citing a Joint Commission statistics page for two of every three sentinel events involving communication failures, which is a different denominator and a broader category. Every figure in the alert is footnoted to a document not read here.
Peer-reviewed#
Weber, D. E., MacGregor, S. C., Provan, D. J. and Rae, A. (2018). 'We can stop work, but then nothing gets done.' Factors that support and hinder a workforce to discontinue work for safety#
Safety Science, 108, 149-160. Safety Science Innovation Lab, Griffith University
- Method
- Qualitative study. Ten focus groups with workers in a range of roles in the liquefied petroleum gas industry, examining an explicit organisational Authority to Stop an Unsafe Task.
- Finding
- Stopping work for safety was reported as challenging at the sharp operational end despite the authority existing and carrying no formal penalty. The authors conclude that stopping an unsafe task 'does not solely hinge on the willingness of individual workers to stop, but also depends on contextual factors surrounding the stop work decision'. A participant's account supplies the title: the authority exists, using it produces no drama, and then nothing gets done, so the work resumes as before.
- What it supports
- That granting an authority is not the same as making it usable, and that the decisive factor is what happens to the work and the worker afterwards rather than the existence of the permission.
- What it does not support
- Any rate or frequency. Ten focus groups in one industry in one country, qualitative by design, with no measurement of how often stops occurred or should have.
Peer-reviewed#
Dauth, W., Findeisen, S., Suedekum, J. and Woessner, N. (2021). The Adjustment of Labor Markets to Robots#
Journal of the European Economic Association, 19(6), 3104-3153
- Method
- German administrative worker and plant data, 1994 to 2014, with a shift-share instrument for robot exposure.
- Finding
- Incumbent workers largely kept their jobs and moved into new, higher-quality tasks within their original plants. The cost fell instead on young labour-market entrants, who shifted away from vocational manufacturing training towards university.
- What it supports
- That the damage from automation falls on skill FORMATION rather than skill possession. Twenty years of German manufacturing data making the missing-rungs argument before anyone applied it to knowledge work.
- What it does not support
- That generative AI will behave like industrial robots. The technologies and the tasks differ substantially.
Peer-reviewed#
Wong, A., Otles, E., Donnelly, J. P. et al. (2021). External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients#
JAMA Internal Medicine, 181(8), 1065-1070. DOI 10.1001/jamainternmed.2021.2626
- Method
- Retrospective external validation cohort study. 27,697 patients aged 18 or over across 38,455 hospitalisations at Michigan Medicine, 6 December 2018 to 20 October 2019. Sepsis occurred in 7 per cent of hospitalisations.
- Finding
- The Epic Sepsis Model achieved a hospitalisation-level area under the curve of 0.63 (95% CI 0.62-0.64), against the 0.76-0.83 cited by its developer. At the alerting threshold in clinical use, sensitivity was 33 per cent, specificity 83 per cent, positive predictive value 12 per cent. It did not identify 1,709 of 2,552 septic hospitalisations (67 per cent), 60 per cent of whom received timely antibiotics anyway, while crossing the alert threshold in 18 per cent of all hospitalisations (6,971 of 38,455), requiring eight patients to be evaluated per case of sepsis found.
- What it supports
- That a proprietary clinical prediction model deployed at national scale can perform far below its developer's stated figures when validated independently, and that nobody had checked. The authors' own conclusion is that widespread adoption despite poor performance raises fundamental concerns about sepsis management nationally.
- What it does not support
- That all clinical prediction models fail, or that this model performs identically elsewhere. It is one model at one academic health system, and the authors note the theoretical alert burden does not account for real-world trigger criteria and lockouts.
Peer-reviewed#
Kanazawa, K., Kawaguchi, D., Shigeoka, H. and Watanabe, Y. (2022). AI, Skill, and Productivity: The Case of Taxi Drivers#
NBER Working Paper 30612; published in Management Science, 72(2), 1376-1388 (2026)
- Method
- Driver-level data from a Japanese taxi fleet through the rollout of an AI demand-prediction system.
- Finding
- Productivity gains accrued almost entirely to LOW-skilled drivers, narrowing the gap between best and worst by 14 percent.
- What it supports
- That the novice-boost pattern found in customer support and software also appears in manual frontline work.
- What it does not support
- Included deliberately as a disconfirming case for a tidy story. Anyone arguing that AI levels up white-collar workers while degrading frontline ones has to explain this result, which runs the other way.
Peer-reviewed#
Kesavan, S., Lambert, S. J., Williams, J. C. and Pendem, P. K. (2022). Doing Well by Doing Good: Improving Retail Store Performance with Responsible Scheduling Practices at the Gap, Inc.#
Management Science, 68(11), 7818-7836
- Method
- Randomised field experiment across 28 Gap stores in San Francisco and Chicago over nine months, November 2015 to August 2016, analysed as intent-to-treat.
- Finding
- Restoring schedule predictability and worker control raised productivity 5.1 percent, with sales up 3.3 percent and labour hours down 1.8 percent.
- What it supports
- That algorithmic optimisation of scheduling is not merely harsh but operationally counterproductive: giving humans back control improved the numbers the algorithm was optimising.
- What it does not support
- Anything directly about generative AI. It concerns algorithmic scheduling, which is an older and different technology.
Peer-reviewed#
Lang, K., Josefsson, V., Larsson, A.-M. et al. (2023). Artificial intelligence-supported screen reading versus standard double reading in the Mammography Screening with Artificial Intelligence trial (MASAI): a clinical safety analysis#
The Lancet Oncology, 24(8), 936-944. DOI 10.1016/S1470-2045(23)00298-X
- Method
- Planned interim safety analysis of a randomised, controlled, non-inferiority, single-blinded screening accuracy trial. 80,033 women aged 40-80 screened at four sites in southwest Sweden between April 2021 and July 2022, randomised 1:1 to AI-supported screen reading or standard double reading by two radiologists.
- Finding
- Cancer detection was six per 1,000 screened women with AI support against five per 1,000 with standard double reading, 41 more cancers detected. The false-positive rate was 1.5 per cent in both arms. Screen readings fell from 83,231 in the control arm to 46,345 in the AI arm, a 44 per cent reduction in screen-reading workload.
- What it supports
- That a triage-plus-detection-support workflow with a radiologist retaining the recall decision can hold detection while roughly halving reading volume, in a randomised population-based programme.
- What it does not support
- Patient benefit, which the interim analysis was not designed to test. It is also one mammography device, one AI system, one country, and moderately to highly experienced readers, which the authors state as limits on generalisability.
Peer-reviewed#
Goh, E., Gallo, R., Hom, J. et al. (2024). Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial#
JAMA Network Open, 7(10), e2440969. DOI 10.1001/jamanetworkopen.2024.40969
- Method
- Single-blind randomised clinical trial, 29 November to 29 December 2023. 50 US-licensed physicians (26 attendings, 24 residents) in family medicine, internal medicine or emergency medicine, randomised to GPT-4 or to conventional resources, working through up to six clinical vignettes, graded blind against a validated diagnostic reasoning rubric. 244 cases completed.
- Finding
- Median diagnostic reasoning score per case was 76 per cent (IQR 66-87) with the LLM and 74 per cent (IQR 63-84) with conventional resources: an adjusted difference of 2 percentage points (95% CI -4 to 8, p=0.60). Median time per case was 519 seconds against 565, adjusted difference -82 seconds (95% CI -195 to 31, p=0.20). The LLM alone scored a median 92 per cent, 16 percentage points above the conventional-resources group (95% CI 2-30, p=0.03).
- What it supports
- That adding a capable model to a physician did not, in this trial, improve diagnostic reasoning or save time, while the same model working alone outperformed both groups of clinicians. The performance of a human-plus-model pairing cannot be inferred from the performance of either part.
- What it does not support
- That LLMs should diagnose autonomously, which the authors explicitly reject. Six curated vignettes exclude history-taking, examination, context and time, which is most of clinical reasoning. Participants received no prompt-engineering training, and the authors offer prompting and interaction design as explanations.
Working paper#
Lee, Y. S., Iizuka, T. and Eggleston, K. (2024). Robots and Labor in Nursing Homes#
NBER Working Paper 33116
- Method
- Original facility-level panel of Japanese nursing homes, using regional robot subsidies as an instrument for adoption.
- Finding
- Robot adoption RAISED employment and improved retention, most strongly for non-regular staff, reallocated worker effort towards direct care, and improved quality: less use of physical restraint and fewer pressure ulcers.
- What it supports
- That automation can absorb routine physical work and upgrade the human job rather than hollow it. The cleanest counter-case in this evidence base, drawn from care work rather than knowledge work.
- What it does not support
- That this generalises. Japanese long-term care faces acute labour shortage, so robots substituted for vacancies rather than for people, which is a very particular condition.
Peer-reviewed#
Patel, V. R., Liu, M., Worsham, C. M. and Jena, A. B. (2024). Alzheimer's disease mortality among taxi and ambulance drivers: population based cross sectional study#
The BMJ, 387, Christmas issue, 17 December 2024. DOI 10.1136/bmj-2024-082194. PROVENANCE: bmj.com could not be reached from either fetcher on 5 September 2026, so the figures below are taken from BMJ Group's own press release for the paper and the Science Media Centre briefing, both read at source that day, rather than from the full text. Re-read and confirm at bmj.com when access allows.
- Method
- Population-based cross-sectional study of US death certificates from the National Vital Statistics System, 1 January 2020 to 31 December 2022, covering 443 occupations and nearly 9 million deaths with occupational information. Usual occupation is the one in which the decedent spent most of their working life. Adjusted for age at death and sociodemographic factors. Published in the BMJ's Christmas issue, which is peer reviewed and deliberately light in subject matter.
- Finding
- 3.9 per cent of deaths (348,328) had Alzheimer's disease listed as a cause. Among 16,658 taxi drivers, 171 did (1.03 per cent); among 1,348 ambulance drivers, 10 did (0.74 per cent). After adjustment, taxi and ambulance drivers had the lowest proportion of any occupation examined (1.03 and 0.91 per cent) against 1.69 per cent for the general population. The pattern did not appear among bus drivers or pilots, who follow predetermined routes, nor for other dementias.
- What it supports
- That an occupational association exists at national scale, and that it is specific to route-generating rather than route-following driving. The authors' own summary is the right one: 'We view these findings not as conclusive, but as hypothesis generating.'
- What it does not support
- That navigation protects anyone. Three limitations do most of the damage, two of them raised by Tara Spires-Jones for the Science Media Centre. The drivers died at around 64 to 67 against 74 for other occupations, and Alzheimer's onset is typically after 65, so some may not have lived long enough to develop it. Women were 10 to 22 per cent of the drivers against 48 per cent elsewhere, on a disease women are more likely to develop. And no brain imaging was involved, so the hippocampal mechanism is hypothesis rather than measurement. Robert Howard, reviewing it alongside her, judged it premature to suggest drivers turn off their satnavs to prevent dementia.
Peer-reviewed#
Yu, F., Moehring, A., Banerjee, O., Salz, T., Agarwal, N. and Rajpurkar, P. (2024). Heterogeneity and predictors of the effects of AI assistance on radiologists#
Nature Medicine, 30(3), 837-849
- Method
- 140 radiologists, 15 chest X-ray diagnostic tasks, roughly 5,190 observations, randomised AI assistance, with empirical-Bayes shrinkage to separate genuine individual differences from noise.
- Finding
- The effect of AI assistance diverged sharply between radiologists, from strongly positive to strongly negative. Experience, subspecialty and prior familiarity with AI all failed to predict who would benefit, and lower performers did not consistently gain.
- What it supports
- That the effect of AI assistance on expert performance is individual and currently unpredictable, so a policy of giving everyone the tool will help some professionals and harm others with no way to tell in advance which.
- What it does not support
- That AI assistance is bad on average, or that the pattern holds outside diagnostic imaging.
Peer-reviewed#
Budzyn, K., Roman'czyk, M., Kitala, D. et al. (2025). Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study#
The Lancet Gastroenterology and Hepatology, 10(10), 896-903. DOI 10.1016/S2468-1253(25)00133-5
- Method
- Retrospective observational study nested in the ACCEPT trial, four Polish endoscopy centres. 1,443 colonoscopies performed WITHOUT AI assistance (795 before and 648 after AI was introduced) by 19 endoscopists averaging 27.6 years of experience, ranging from 8 to 39 years.
- Finding
- Adenoma detection rate in unassisted colonoscopy fell from 28.4 percent before AI exposure to 22.4 percent after, a drop of 6.0 percentage points (p=0.0089; adjusted odds ratio 0.69).
- What it supports
- Measured deskilling in highly experienced professionals, in unassisted performance, within months of routine AI exposure. The strongest direct evidence that capability degrades when a tool takes over the judgement, rather than merely a plausible mechanism.
- What it does not support
- Causation with certainty: it is observational, not randomised, and other changes over the period cannot be fully excluded. It is also one procedure in one country, and detection rate is a proxy for skill rather than skill itself.
Peer-reviewed#
Heinz, M. V., Mackin, D. M., Trudeau, B. M. et al. (2025). Randomized Trial of a Generative AI Chatbot for Mental Health Treatment#
NEJM AI, 2(4). DOI 10.1056/AIoa2400802
- Method
- Randomised controlled trial, N=210 adults recruited by national advertising, with major depressive disorder, generalised anxiety disorder or clinically high risk for feeding and eating disorders. Four-week intervention with Therabot, an expert-fine-tuned generative chatbot built on a hand-curated CBT corpus, against a WAITLIST control, with four weeks of follow-up.
- Finding
- Statistically significant symptom reductions across all three diagnostic groups at four and eight weeks. 95 per cent of participants engaged, averaging 260 messages and 6.18 hours over four weeks. Therapeutic alliance was comparable to outpatient psychotherapy (mean WAI 3.59). Expressions of suicidal ideation required staff intervention 15 times, and inappropriate responses required correction 13 times.
- What it supports
- That a purpose-built, expert-curated generative chatbot can produce measurable symptom improvement over four weeks, and that users form a working alliance with it.
- What it does not support
- That it is safe unsupervised, or that the effect is the chatbot rather than attention and expectation. The control was a waitlist rather than an active comparator, so the treatment effect cannot be separated from the effect of receiving something. The developer confirmed to the FDA advisory committee that humans reviewed all messages in near real time and conducted risk assessments where needed, so the safety record describes a supervised system rather than an autonomous one.
Peer-reviewed#
Lukac, P. J., Turner, W., Vangala, S., Chin, A. T., Khalili, J., Shih, Y.-C. T., Sarkisian, C., Cheng, E. M. and Mafi, J. N. (2025). Ambient AI Scribes in Clinical Practice: A Randomized Trial#
NEJM AI, 2(12). DOI 10.1056/AIoa2501000. Preprint: medRxiv 2025.07.10.25331333
- Method
- Parallel three-arm pragmatic randomised clinical trial at one US health system. 238 outpatient physicians across 14 specialties randomised 1:1:1 by covariate-constrained randomisation to Microsoft DAX, Nabla or usual care, 4 November 2024 to 3 January 2025, with the second intervention month compared to baseline.
- Finding
- Time writing a note fell by an estimated 18 seconds in the control arm, 23 seconds in the DAX arm and 41 seconds in the Nabla arm. Only Nabla differed significantly from control (-9.5 per cent, 95% CI -17.2 to -1.8, p=0.02); DAX did not (-1.7 per cent, 95% CI -9.4 to +5.9, p=0.66). Scribe users improved on Mini-Z burnout (+2.76, p<0.001), task load (-35.8, p=0.01) and work exhaustion (-0.27, p=0.01), with no significant difference between the two products on any psychometric. Roughly 15 per cent of physicians given a tool never used it.
- What it supports
- That ambient documentation produces a small, product-dependent time saving and a more consistent wellbeing effect, measured against a randomised control rather than against a before-and-after.
- What it does not support
- A general time saving for ambient AI. One health system, majority female sample, a two-month contract-limited window, and the authors flag that the electronic record's own time metrics do not count editing done inside the scribe platform, so reported savings may be overstated. They note this limitation affects all studies using those metrics.
Peer-reviewed#
Moore, J., Grabb, D., Agnew, W., Klyman, K., Chancellor, S., Ong, D. C. and Haber, N. (2025). Expressing stigma and inappropriate responses prevents LLMs from safely replacing mental health providers#
Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency (FAccT). DOI 10.1145/3715275.3732039
- Method
- Mapping review of therapy guides used by major medical institutions to identify the requirements of a therapeutic relationship, followed by several experiments testing current models, including gpt-4o, against those requirements in naturalistic therapy settings.
- Finding
- Models expressed stigma towards people with mental health conditions and responded inappropriately to common and critical presentations, including encouraging delusional thinking, which the authors attribute to sycophancy. The pattern persisted in larger and newer models.
- What it supports
- That the failures are structural rather than a matter of model size, and that current safety training does not remove them.
- What it does not support
- How often this occurs in real use, or what happens with purpose-built clinical systems rather than general-purpose models. These are constructed scenarios, not observed patient interactions.
Peer-reviewed#
Nilsson, A. et al. (Karolinska Institutet) (2025). Algorithmic management is associated with psychological distress, musculoskeletal pain, and occupational accidents: a cross-sectional study in logistics#
International Archives of Occupational and Environmental Health, 98
- Method
- Survey of Swedish logistics workers, February to July 2024, 978 respondents (592 drivers, 378 warehouse), using an eleven-item algorithmic-management exposure scale with adjusted models.
- Finding
- Higher exposure to algorithmic management was associated with greater psychological distress, more occupational accidents and more musculoskeletal pain.
- What it supports
- That where AI meets frontline work the output is measured in bodies rather than in output quality, a category almost entirely absent from white-collar productivity research.
- What it does not support
- Causation: it is cross-sectional and self-reported. Workers under strain may also perceive management as more algorithmic.
Qualitative field study#
Ehsan, U., Passi, S., Saha, K., McNutt, T., Riedl, M.O. and Alcorn, S. (2026). From Future of Work to Future of Workers: Addressing Asymptomatic AI Harms to Foster Dignified Human-AI Interaction#
Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI '26), ACM. DOI 10.1145/3772318.3791081. Preprint arXiv:2601.21920, 29 January 2026
- Method
- Twelve months of situated fieldwork through the first year of routine use of a commercial AI-assisted radiotherapy treatment-planning system across a five-site North American hospital group. 42 participants: 15 radiation oncologists, 12 medical physicists, 7 dosimetrists, 8 administrators, with 2 to 24 years of experience. 52 think-aloud sessions, 24 interviews staged across months 2 to 11, five participatory workshops, 63 hours transcribed, analysed by grounded theory.
- Finding
- Measured operational gains and reported capability loss ran together. Planning cycles shortened by roughly 15 per cent and confidence rose, while by month nine several dosimetrists said their unaided proficiency had worsened over the year, one saying they had grown slower without the tool. The authors name two things the estate has nowhere else: intuition rust, the gradual dulling of expert judgement beneath intact output, and identity commoditisation, the erosion of professional dignity as practitioners describe becoming AI babysitters, button-pushers and bystanders in their own practice. They call these harms asymptomatic because the organisation's own measures showed only the improvement.
- What it supports
- That the instruments an organisation uses to judge an AI deployment can register the gain and be structurally blind to the cost, and that practitioners perceive the loss long before any dashboard does. It is the best available account of what deskilling feels like from inside a high-stakes clinical specialty, and the only source here on occupational identity.
- What it does not support
- Any measured deskilling. The skill claims are self-reported, plus one informal unaided exercise with a couple of dosimetrists during a workshop. There is no controlled comparison, no pre and post measurement and no quantified skill outcome, so it cannot be set beside Budzyn as a second measured result. IMPORTANT, this study is widely misdescribed in secondary summaries as radiologists using generative AI. It is neither. The participants plan radiotherapy treatment rather than interpret images, and the system is optimisation-based, not generative. Do not repeat that description.
Peer-reviewed#
Gommers, J., Lang, K., Hofvind, S. et al. (2026). Interval cancer, sensitivity, and specificity comparing AI-supported mammography screening with standard double reading without AI in the MASAI study#
The Lancet. DOI 10.1016/S0140-6736(25)02464-X. Published 29 January 2026
- Method
- Full results of a randomised, controlled, non-inferiority, single-blinded, population-based screening-accuracy trial. Over 100,000 women screened at four Swedish sites between April 2021 and December 2022, with two years of follow-up.
- Finding
- Interval cancers fell from 1.76 per 1,000 women (93/52,872) in the control arm to 1.55 per 1,000 (82/53,043) in the AI arm, a 12 per cent reduction. There were 16 per cent fewer invasive (75 v 89), 21 per cent fewer large (38 v 48) and 27 per cent fewer aggressive-subtype (43 v 59) interval cancers. Cancers detected at screening rose from 74 per cent (262/355) to 81 per cent (338/420) of all cases. False positives were 1.5 per cent in the intervention arm and 1.4 per cent in the control arm.
- What it supports
- That AI-supported screen reading, with a radiologist retaining the recall decision, reduced interval cancers over two years in a randomised population-based programme. The only outcome-level randomised evidence of this kind in clinical AI.
- What it does not support
- Generalisation beyond one country, one mammography device, one AI system and experienced readers, all stated as limitations by the authors. Mortality was not an endpoint, and cost-effectiveness was not assessed. The first author's own summary states the design does not support replacing radiologists, since at least one still reads every case.
The international picture
Evidence from outside the US, UK and Nordic economies, including sources published in other languages.
Institutional modelling#
Korea Development Institute (KDI) (2023). Changes in the labour market due to artificial intelligence and policy directions (Research Report 2023-03)#
KDI, published in Korean
- Method
- Expert and GPT-4 capability ratings applied to Korean occupational profiles, a KDI survey of 800 firms in September 2023, and firm-panel econometrics.
- Finding
- 38.8 percent of jobs are technically automatable across more than 70 percent of their tasks, yet only 2.7 percent of firms with ten or more staff had adopted AI. Realised effects showed no aggregate employment change, lower earnings, and the impact concentrated on YOUNGER, tertiary-educated workers and women.
- What it supports
- That the gap between technical potential and actual adoption is enormous, and that where effects appear they fall on the young and educated rather than the low-skilled.
- What it does not support
- That the Korean pattern transfers. Korea has unusually high tertiary education rates and a distinctive labour market.
Institutional survey#
Saudi Data and AI Authority (SDAIA) (2023). AI Ethics Principles, Version 1.0#
SDAIA, September 2023
- Method
- National AI ethics guidance, seven principles with an assessment checklist annexe covering the AI system lifecycle. Non-binding guidance rather than statute. Version 1.0 read in full at source; whether a later version exists could not be checked because SDAIA's main host rejects automated requests.
- Finding
- The checklist annexe puts this question to designers at the plan and design stage: "Does your AI system design prevent overconfidence in or overreliance on the AI system with necessary human intervention mechanisms?" It also asks whether human oversight processes carry defined KPIs and assigned responsibility. The operative text states that decisions which are irreversible or life-and-death "should trigger human oversight and final determination", and rules out social scoring and mass surveillance.
- What it supports
- That over-reliance on AI has been named as a design defect to be engineered against in a national governance instrument. On the evidence gathered here it is the only Gulf instrument that does so.
- What it does not support
- Any obligation. It is guidance, not law, the over-reliance language sits in a checklist annexe rather than in the principle text, and Saudi Arabia had no binding AI statute as of August 2026, only a draft responsible AI policy out for consultation from 2 April 2026. Version 1.0 dates from September 2023 and may have been superseded.
Institutional survey#
Government of the United Arab Emirates (2024). The UAE Charter for the Development and Use of Artificial Intelligence#
UAE Legislation portal, issued 10 June 2024
- Method
- National charter of twelve principles. Statement of principle with no duty-holder, no enforcement mechanism and no competence requirement.
- Finding
- Principle 6, Human Oversight, "emphasizes the irreplaceable value of human judgment and human oversight over AI, aligning with ethical values and social standards to correct any errors or biases that may arise". The charter is silent on deskilling, over-reliance and any obligation to train or assess the humans doing the overseeing.
- What it supports
- That human oversight is stated as a national principle in the UAE.
- What it does not support
- That it is operational. There is no named duty-holder, no enforcement, no competence standard and no test of whether oversight is real. The UAE had no federal AI statute as of August 2026.
Institutional modelling#
Institut fur Arbeitsmarkt- und Berufsforschung (IAB), Germany (2024). Folgen des technologischen Wandels fur den Arbeitsmarkt (Consequences of technological change for the labour market: it is above all the highly qualified who feel digitalisation)#
IAB-Kurzbericht 5/2024, published in German
- Method
- The Substituierbarkeitspotenziale series, 2022 wave. Three independent coders score more than 9,000 tasks in the BERUFENET expert database across roughly 4,600 occupations.
- Finding
- Substitutability rose about ten percentage points for degree-level expert occupations between 2019 and 2022, and was roughly flat for helper occupations. IAB frames AI as relief for skills shortages rather than as displacement.
- What it supports
- That the German expert assessment puts the pressure on the highly qualified, which inverts the assumption that automation threatens the least skilled first.
- What it does not support
- Realised outcomes. Substitutability is technical potential assessed by coders, not what employers did.
Institutional survey#
LaborIA (French Ministry of Labour, Inria and Matrice) (2024). Etude des impacts de l'IA sur le travail: Rapport d'enquete LaborIA Explorer (Study of the impacts of AI on work)#
LaborIA, published in French
- Method
- Telephone survey of 250 decision-makers in firms with more than 50 staff (42 with AI deployed), longitudinal interviews with 10 decision-makers across three waves, and six ethnographic field sites.
- Finding
- Names a conflit de rationalite, a clash of rationalities: managers justify AI by error reduction (81 percent), performance (75 percent) and removing drudgery (74 percent), while fieldwork shows workers becoming the system's de facto trainers.
- What it supports
- That what management believes AI is doing and what workers experience it doing can diverge systematically inside the same organisation. A work-psychology and ergonomics frame rather than task-exposure modelling.
- What it does not support
- Scale. The qualitative core rests on six sites and ten repeated interviews.
Institutional survey#
BAuA, ZEW, IAB and BIBB, Germany (2025). Digitalisierung und Wandel der Beschaftigung, DiWaBe 2.0 (Digitalisation and the transformation of employment)#
Bundesanstalt fur Arbeitsschutz und Arbeitsmedizin, published in German
- Method
- Representative 2024 survey of roughly 9,800 employees subject to social insurance, linkable to administrative employer and employee records.
- Finding
- More than half already use AI at work but largely informally. Use ranges from about a third of unqualified workers to around 80 percent of those with a degree or Meister qualification. There was NO difference in training participation between AI users and non-users.
- What it supports
- That AI use is spreading through workplaces without any corresponding increase in training, which is the adoption-without-redesign pattern measured directly at national scale.
- What it does not support
- What that absence of training does to capability over time. The survey is a single 2024 snapshot.
Institutional survey#
Bundesinstitut fuer Berufsbildung (2025). Angespannte Lage auf dem Ausbildungsmarkt (Tense situation in the training market)#
BIBB press release 40/2025, 10 December 2025
- Method
- Official German statistics on the dual vocational training system, counting newly concluded training contracts, unfilled training places and unplaced applicants as at 30 September 2025.
- Finding
- Around 476,000 new dual training contracts in 2025, down 2.1 per cent on 2024 and the second consecutive annual fall. 54,400 training places went unfilled. At the same time around 84,400 young people had not found a training place, up 19.9 per cent and the highest since 2010, with the share of unsuccessful applicants at 15.1 per cent, the highest since the end of the 2009 financial crisis.
- What it supports
- That the world's most admired apprenticeship system is failing to match young people to places in both directions at once: record unplaced applicants alongside tens of thousands of empty places. The bridge is not merely narrowing, it has stopped connecting.
- What it does not support
- That AI caused it. BIBB attributes the tension to matching problems, regional and occupational mismatch and economic conditions, and makes no AI claim.
Institutional survey#
General Authority for Statistics (GASTAT), Saudi Arabia (2025). Establishments ICT Access and Usage Statistics 2025#
GASTAT, 2025
- Method
- National statistical survey of establishments, methodology stated as aligned with UNCTAD international standards. Enterprise-side only.
- Finding
- 33.1 per cent of establishments use artificial intelligence technologies, a growth of 20.0 per cent against 2024. By sector: information and communication 61.1 per cent, financial and insurance 52.9 per cent, education 51.0 per cent, manufacturing 30.7 per cent, wholesale and retail 30.2 per cent. The companion household and individuals survey for the same year contains no AI indicator at all.
- What it supports
- That one Gulf state measures enterprise AI adoption with a published, standards-aligned method. It is the only official national AI adoption statistic found across Saudi Arabia, the UAE and Qatar.
- What it does not support
- Anything about citizens, or about employment. Saudi Arabia measures enterprise adoption but not individual use, and publishes no AI employment series. Adoption is also self-reported use of a technology category, not a measure of capability or of value obtained.
Institutional survey#
Japan Institute for Labour Policy and Training (JILPT) (2025). Survey on the impact of workplace AI adoption on working styles (Research Series No. 256)#
JILPT, published in Japanese, designed with the OECD
- Method
- 22,000 employees, stratified on 2020 Census occupation, employment type, sex and age. Fieldwork May to June 2024.
- Finding
- Only 12.9 percent report any firm AI use and 8.4 percent use it themselves. Among users, reports of improved job quality and wellbeing outweighed reports of decline, and the gain was markedly larger where the employer had consulted staff and funded training.
- What it supports
- That the effect of AI on how work feels is conditional on how the employer introduced it, not determined by the technology. Direct empirical support for the design-rather-than-drift argument, from a 22,000-person sample.
- What it does not support
- Long-run capability effects, and Japanese adoption rates are far below US levels so the user group is early and unusual.
Institutional modelling#
Lu, Y. and Gui, L. (Bulletin of the Chinese Academy of Sciences) (2025). Analysis of the impact of artificial intelligence technology on employment and income in China#
Bulletin of the Chinese Academy of Sciences, 40(4), 642-651, published in Chinese
- Method
- Policy synthesis of the Chinese empirical literature and official statistics, under National Social Science Fund major project 23ZDA100.
- Finding
- Between 2018 and 2023 the substitution effect outweighed complementarity: a one percent rise in industrial robots reduced firm labour demand by 0.18 percent. Roughly 200 million people, 27 percent of employment, are in flexible work with 37 percent social-insurance coverage.
- What it supports
- That in the largest manufacturing economy the measured balance so far has been substitution rather than augmentation, and the policy response centres on social security redesign rather than retraining.
- What it does not support
- Comparability with service-sector generative AI in high-income economies. This is largely industrial robotics.
Operator account#
Moquim, S. A., Vice President, Saudi Data and Artificial Intelligence Authority, interviewed by Saeed al-Abyad (2025). SDAIA: Saudi AI Platform Baseer Boosts Crowd, Security Control During Hajj#
Asharq Al-Awsat, dateline Jeddah, 4 June 2025; read at source 4 September 2026
- Method
- On-the-record interview with the operating authority's vice president. Descriptive throughout, with no evaluation and no data.
- Finding
- SDAIA operates Baseer, built with the Ministry of Interior, using AI algorithms and computer vision on live feeds to detect crowd density and distribution within the Grand Mosque and to pinpoint overcrowded zones such as the Tawaf area moment by moment, stated as enabling authorities to act swiftly to prevent overcrowding or stampedes. Companion platforms Sawaher and Sawaher Qiyada analyse live security camera feeds. The Smart Makkah Operations Center coordinates them, and biometric systems run at 12 international airports across 8 countries under the Makkah Route initiative.
- What it supports
- That the same public authority which published Saudi Arabia's AI ethics principles, including the requirement that irreversible or life-and-death decisions trigger human oversight and final determination, also builds and operates the system to which that requirement applies.
- What it does not support
- That the oversight works, or that anyone has tested it. Nothing here measures how often a commander overrides the system, what happens when the prediction is wrong, or whether unaided crowd-reading skill is maintained. The speaker heads the authority whose systems he is describing.
Institutional survey#
National Centre for Vocational Education Research (2025). Apprentices and trainees 2025#
NCVER, Australia, March and June quarter releases 2025
- Method
- Official Australian administrative counts of apprentice and trainee training contracts.
- Finding
- 320,830 in-training contracts at 31 March 2025, down 7.9 per cent year on year, with trade down 3.2 per cent and non-trade down 17.9 per cent. By the June quarter the fall was 11.3 per cent, with non-trade down 20.2 per cent. NCVER attributes the non-trade decline partly to the conclusion of the Boosting Apprenticeship Commencements subsidy.
- What it supports
- That entry-level training volume responds sharply and quickly to subsidy withdrawal, and that the response is concentrated in the non-trade occupations closest to office work.
- What it does not support
- An AI effect. NCVER names policy change, and the timing follows the 2022 subsidy withdrawal rather than any AI milestone.
Institutional modelling#
Azim Premji University (2026). State of Working India 2026: Youth in the Labour Market#
Azim Premji University, Bengaluru
- Method
- Analysis of National Sample Survey employment data 1983 to 2011, all quarters of the Periodic Labour Force Survey 2017 to 2024, plus AISHE, NCVT-MIS and CMIE-CPHS.
- Finding
- Roughly five million graduates enter the Indian labour market each year against about 2.8 million finding work, with graduate unemployment near 40 percent for 15 to 25 year olds. The report explicitly declines to attribute this to AI.
- What it supports
- That in the world's most populous labour market the early-career crisis is a demand-side bottleneck that long predates AI. A necessary corrective to reading every graduate hiring problem as an AI story.
- What it does not support
- Anything about AI's effect in India, which it deliberately does not claim.
Compiled review#
Cedefop (2026). Call for abstracts: apprenticeships in the age of AI#
Cedefop, European Centre for the Development of Vocational Training, 5 August 2026
- Method
- Framing statement for the 2027 joint Cedefop and OECD apprenticeship symposium. Position paper, not a study.
- Finding
- Cedefop states: "AI takes over baseline tasks which were in many cases performed by apprentices or apprenticeship graduates who get entry-level roles once their programmes are completed. As apprenticeships continue to expand into new fields and occupations, they may also be exposed to a decline in entry-level openings." It adds that the contraction appears concentrated in white-collar roles, so the effect is likely to be uneven and occupation-specific rather than uniform.
- What it supports
- That a European Union agency has now named the mechanism this research describes, in its own words, and applied it specifically to apprenticeships. It is the clearest institutional statement that the training route and the entry-level job are the same thing.
- What it does not support
- Anything measured. It is a call for papers, which means Cedefop is asking the question rather than answering it, and it says the effect is uneven.
Institutional survey#
Census and Statistics Department, Hong Kong SAR (2026). Report on the Survey on Information Technology Usage and Penetration in the Business Sector, 2025 Edition#
Census and Statistics Department, Hong Kong, released 27 February 2026
- Method
- National business survey, fieldwork March to December 2025. Full report and all thirty tables read at source, together with three companion household and ICT publications.
- Finding
- Artificial intelligence appears nowhere in the report: not in the tables, not in the explanatory notes, not in the definitions. The survey's ICT categories are cloud computing at 98.1 per cent, QR codes at 37.6 per cent, RFID at 20.3 per cent, internet of things at 7.4 per cent and augmented or virtual reality at 1.5 per cent. Hong Kong's statistical office measures AR and VR adoption and does not ask about AI at all. The 2023 edition also had no AI category, and no plan to add one has been announced.
- What it supports
- That Hong Kong publishes no official statistic on AI adoption or AI-related employment, which means every circulating Hong Kong AI adoption figure comes from a non-government survey with a self-selected sample.
- What it does not support
- That AI adoption in Hong Kong is low. It establishes that it is unmeasured by the government, which is a different and in some ways more useful fact.
Institutional modelling#
INSEE, France (2026). Note de conjoncture: digital investment, artificial intelligence and youth employment#
Institut national de la statistique et des etudes economiques, March 2026, published in French
- Method
- National accounts and quarterly employment files, with an error-correction model estimated 1990Q1 to 2019Q4.
- Finding
- French employment of 15 to 29 year olds, excluding apprentices, fell 7.4 percent year on year in IT services, 5.8 percent in publishing and 3.7 percent in management consulting in Q4 2025, against minus 0.7 percent across the market sector overall.
- What it supports
- A European national-statistics office finding the same entry-level pattern reported in the US, which makes the signal considerably harder to dismiss as an American artefact.
- What it does not support
- That AI caused it. INSEE explicitly cautions against attributing the fall to AI alone.
Institutional survey#
International Labour Organization (2026). Global Employment Trends for Youth 2026: Back to the future#
ILO, 11 August 2026
- Method
- ILO global modelled estimates of youth employment, unemployment and NEET status.
- Finding
- Global youth unemployment 12.4 per cent in 2025, 67 million people aged 15 to 24. The NEET rate is 20 per cent, over 257 million. Youth unemployment rose in 8 of 11 subregions between 2023 and 2025, with Northern America rising from 8.3 to 9.8 per cent. The ILO estimates 6.1 per cent of jobs held by workers aged 15 to 29 are in occupations highly exposed to AI-related change.
- What it supports
- That youth labour market conditions deteriorated across most of the world between 2023 and 2025, which is the context any national apprenticeship policy is operating in.
- What it does not support
- That AI caused it. The 6.1 per cent exposure figure is an occupational overlap measure, not a measured displacement.
Institutional modelling#
Jung, J. and Katz, R. (CEPAL / ECLAC) (2026). Impacto economico de la inteligencia artificial en America Latina (The economic impact of artificial intelligence in Latin America)#
United Nations Economic Commission for Latin America and the Caribbean, published in Spanish
- Method
- Theoretical and econometric modelling of AI's macroeconomic effect through skilled-labour productivity across the region.
- Finding
- The gains run through skilled labour, and the binding constraint across Latin America is human-capital formation and low investment. The regional risk is UNDER-adoption rather than displacement.
- What it supports
- That the framing dominant in rich economies, where the worry is AI doing too much, inverts in middle-income economies, where the worry is that it will not arrive at all.
- What it does not support
- Firm-level outcomes; it is a macro model.
Institutional survey#
Manpower Research and Statistics Department, Ministry of Manpower, Singapore (2026). Labour Market Report, First Quarter 2026#
Ministry of Manpower, Singapore, released 15 June 2026
- Method
- National firm survey by the labour ministry's statistics department. Note the companion press release omits the AI statistic entirely; it appears only in the full report.
- Finding
- 28.5 per cent of firms adopted AI in 2026, highest in information and communications at 74.1 per cent, professional services at 57.5 per cent and financial and insurance services at 56.4 per cent. Only 6.2 per cent reported AI-related reductions in headcount or hiring, against 18.9 per cent reporting redesign of job functions. The ministry's own reading: "AI is currently having a greater impact on job redesign and work processes than on broad-based job displacement."
- What it supports
- That a government labour ministry measuring this finds job redesign running roughly three times ahead of headcount reduction. It is a useful corrective to displacement forecasting.
- What it does not support
- What happens to capability. Redesign is not neutral: the question this research asks is which tasks the redesign removes, and a firm survey of headcount cannot answer it.
Compiled review#
Ministry of Communications and Information Technology, Qatar (2026). Artificial Intelligence in Qatar: Principles and Guidelines for Ethical Development and Deployment#
MCIT Qatar, undated; read at source 28 August 2026
- Method
- National AI ethics guidance. Read in full at source, but with a provenance defect worth recording: the document carries no publication date, no version number and no reference number, and is not retrievable from the ministry's own website. It states that it "is legally non-binding, and adherence to it is voluntary".
- Finding
- Principle 8, assign ultimate accountability to humans, states that "AI systems should not be able to autonomously make decisions of significant consequence" and should "provide users with the ability to appeal or override decisions that have a substantial impact on individuals or society". The phrase human in the loop appears nowhere in the document, and Principle 7, develop a human-centered approach, concerns cultural values, feedback, diverse teams and accessibility rather than oversight.
- What it supports
- That Qatar requires humans to retain control of consequential decisions, in guidance.
- What it does not support
- Any competence duty. Qatar imposes no obligation that the humans exercising control be trained, assessed or kept current, and the document contains no reference to deskilling, over-reliance or automation bias. Widely circulated dates of May 2024 or 2025 for this document are unsourced: it carries no date at all.
Operator account#
Ministry of Interior, Kingdom of Saudi Arabia (2026). Smart Predictive Technologies Enhance Pilgrim Safety and Crowd Management#
Saudi Press Agency, Makkah, 30 May 2026, 1447 AH Hajj season; read at source 4 September 2026
- Method
- The state news agency reporting the Ministry of Interior's own account of its Hajj operation. No methodology, no figures, no independent evaluation.
- Finding
- The ministry states it implemented predictive analytics to anticipate and mitigate congestion and hazardous conditions before they occurred, using an intelligent framework powered by AI and data analytics to accelerate strategic decision-making and, in its own words, to support field commanders. Oversight is described as running through a digital framework of performance indicators under the supervision of specialised personnel.
- What it supports
- That a state operating one of the largest recurring crowd operations in the world describes AI-driven prediction as an input to command decisions, and says so in the language of supporting rather than replacing the commander.
- What it does not support
- Anything measured. There is no figure for accuracy, no override rate, no comparison against unaided command judgement, and no independent evaluation of whether the oversight functions. It is a state news agency reporting a ministry on its own performance.
Institutional modelling#
Rodriguez-Fernandez, M. (Funcas) (2026). Inteligencia artificial y mercado de trabajo en Espana (Artificial intelligence and the labour market in Spain)#
Funcas working paper, published in Spanish
- Method
- The Felten AI occupational exposure index remapped to Spanish occupational classifications, combined with the Q4 2025 Labour Force Survey.
- Finding
- Spain shows medium-high exposure at 27.4 percent but low automation risk at 5.9 percent, against an OECD average near 12 percent, because of its interpersonal and physical occupational mix.
- What it supports
- That national occupational structure, not technology, determines exposure. An economy weighted towards interpersonal and physical work is structurally less automatable.
- What it does not support
- Outcomes. Exposure indices remain estimates of what could be affected.
What the weight of the evidence supports#
Read together rather than one at a time, the base supports a narrower claim than either the optimists or the pessimists make. It does not show that AI makes people less intelligent. It shows something more specific and harder to dismiss: capability weakens where practice stops, offloading decisions are frequently misjudged, confident machine output reliably suppresses scrutiny, and the design of the tool decides whether a person is taught or carried. The Bastani field experiment is the clearest single result, because the same underlying model produced both the best learning outcome and the worst, and the only variable was whether the interface made the student do the work.
It also supports honest uncertainty about size and speed. Most of the workplace evidence is self-reported or correlational, the strongest long-run analogues come from memory and navigation rather than reasoning, and no study has yet run across the years over which professional judgement is actually formed. The responsible position is that the direction is consistent enough, and the damage slow and invisible enough, to design against now rather than after a decade of proof arrives.
A note on the gap this fills#
There is no shortage of evidence about what AI can do. Stanford's AI Index is excellent and openly published, and it measures the machine. What is missing, and what this base is for, is an open, graded account of what increasingly capable AI does to the humans working alongside it. The consultancies hold fragments of this and sell them. The academic literature holds the studies but not the synthesis. Assembling it in public, with the limitations stated rather than buried, is the contribution.
Cite an individual entry#
Every entry has a permanent anchor. To cite one, use its link directly, for example thesuperskills.com/research/evidence#vaccaro-2024. The whole base is also available as structured data at evidence.json for anyone, or anything, that would rather read it that way.
How this connects to the SuperSkills research#
The evidence is separated from the interpretation throughout this site, and this page is the evidence half. The interpretation lives in the research: AI and human judgement, AI and critical thinking, how humans learn with AI, human and AI decision making, what stays human, capability debt and staying valuable in the age of AI. Where a concept is Rahim Hirji's own, such as capability debt, the missed reps or synthetic seniority, it is marked as his. Where it is established, such as cognitive offloading or automation bias, it is attributed to the researchers who developed it.
About this reference
Compiled and maintained by Rahim Hirji, author of SuperSkills: The Seven Human Skills for the Age of AI (Kogan Page, 2026) and founder of The SuperSkills Intelligence Company. Every URL was fetched and confirmed before publication. Entries are added as significant work appears and revised when a study is corrected, retracted or superseded; the Gerlich correction is carried in its entry for exactly that reason. Corrections and omissions are welcome: if something important is missing, or something here is mischaracterised, please say so.
About this research#
Written by Rahim Hirji, author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company.
Cite this
Hirji, R. (2026). The evidence on AI and human capability. The SuperSkills Intelligence Company. Last reviewed 5 September 2026. thesuperskills.com/research/evidence
