Yes, on narrow tasks, by amounts large enough that arguing about them is pointless. Inside real work the answer scatters, and one careful randomised study measured experienced people going slower. At the level of an economy, nothing has shown up yet. This page keeps those three altitudes apart, because almost every argument about AI and productivity is two people describing different ones.
The answer, in one line
On narrow, well-specified tasks the measured gains are large and repeated: writing tasks completed 40 per cent faster with 18 per cent higher rated quality, a standardised coding task 55.8 per cent faster, customer support resolutions up 15 per cent an hour.
The short answer#
On well-specified tasks with a clear finish line, generative AI produces large measured gains. In continuing work inside systems people already know, the measured effect is uncertain and sometimes negative. Nothing has yet appeared in national productivity statistics.
Where the gains are large and well measured#
Noy and Zhang randomised professionals across mid-level writing tasks. Average time fell 40 per cent and rated output quality rose 18 per cent, with the gap between stronger and weaker writers narrowing. Graded entry.
Peng and colleagues ran a randomised trial of GitHub Copilot on 95 professional programmers implementing an HTTP server in JavaScript. Among those who finished, the assisted group took 71 minutes against 161, a 55.8 per cent reduction, p=0.0017. The authors state plainly that they did not examine code quality. Graded entry.
Brynjolfsson, Li and Raymond studied 5,000 customer support agents in the field. Resolutions per hour rose about 15 per cent overall, 30 per cent for the least experienced and 36 per cent for the lowest skill quintile, while the most skilled showed no significant gain. Graded entry.
Three different task types, three large effects, one consistent pattern underneath: the gains land hardest on the people who were furthest from the standard, which is the compression result that runs through this whole literature.
Where they thin out, and where they reverse#
Dell'Acqua and colleagues gave 758 consultants GPT-4 across two task types. Inside the model's competence the assisted consultants were dramatically better. On a task designed to sit just outside it, they did worse than consultants with no AI at all. The authors call the boundary a jagged technological frontier, and its importance is that nobody can see where it runs from inside the work. Graded entry.
METR's 2025 experiment is the one that broke the pattern. Experienced open-source developers, working on their own mature repositories, were measured 19 per cent slower when permitted to use AI tools. The tasks were real, the repositories were ones they knew well, and the finding survives as the most cited counter-result in the field. Graded entry.
Set the two coding studies beside each other and the shape of the disagreement appears. Peng's task was a fresh server built from nothing against a clear specification. METR's tasks were changes to large codebases their authors had lived in for years. Speed arrives where the work is production. It disappears where the work is understanding.
METR has since unsettled its own number, and said so#
In February 2026 METR reported a second randomised study, 57 developers across 143 repositories and more than 800 tasks, and stated that their data now gives an unreliable signal. Between 30 and 50 per cent of developers declined to submit tasks they did not want to do without AI, which selects the sample. The raw results point the other way, a speedup of -18 per cent for returning developers and -4 per cent for new recruits, and every confidence interval crosses zero. They believe developers are probably more sped up in 2026 than in 2025 and say their own data is only very weak evidence for it. Graded entry.
Two things follow, and both matter more than the number. The original study is not retracted, and its perception finding is untouched. Anyone still quoting "19 per cent slower" as the current state of AI coding tools is quoting early-2025 tooling and a design its own authors have replaced.
The perception gap runs in both directions#
The most quoted line from METR is that developers forecast a 24 per cent speed-up, were measured 19 per cent slower, and after finishing the tasks still estimated AI had made them about 20 per cent faster. It is a striking result and it has hardened into a general law: people overstate what AI does for them.
The law does not hold. In the Copilot trial, participants in both arms estimated a 35 per cent improvement when the measured figure was 55.8 per cent, so they understated it by twenty points. METR's own 2026 survey team, reporting a median self-reported speed change of three times, note that only one study has gathered survey and field-experiment results on the same population and metric, and decline to say how far surveys overstate in general. Graded entry.
So the honest statement is narrower and more useful than the slogan. People are poor estimators of their own throughput, in whichever direction the error happens to fall, and asking them is not a measurement. That is a reason to measure, and not a reason to assume the answer is smaller than they say.
Two government trials, 23,500 licences, and no baseline between them#
The Government Digital Service ran 20,000 Copilot licences across twelve organisations for three months. Average self-reported saving was 26 minutes a day, adoption held around 80 per cent, and 82 per cent said they would not want to go back. The report's own conclusions state that it was not possible to identify how the saved time was spent. Graded entry.
The Department for Work and Pensions evaluated 3,549 licences against a stratified comparison group of non-users, and estimated 19 minutes a day across eight routine tasks, statistically significant across specifications, with job satisfaction up 0.56 points. Its own limitations chapter names the absence of baseline data, post-treatment bias and self-selection towards enthusiasts through first-come first-served allocation, which it says may lead to overestimation. Graded entry.
These are the largest organisational deployments anybody has published on, and neither can tell you whether the organisation produced more. They are careful, useful documents about how people experienced a tool. An estimate of 26 minutes, multiplied by a headcount, is a business case built on a survey question, and both departments are more honest about that than the people quoting them.
The economy has not noticed#
Acemoglu's estimate remains the most careful macroeconomic one available: total factor productivity gains of no more than 0.66 per cent over ten years, revised to under 0.53 per cent once the difficulty of hard-to-learn tasks is accounted for. Graded entry.
A gap this wide between task-level results and aggregate ones is not itself surprising, and it has a long history in economics. It does mean that anyone extrapolating from a 40 per cent writing gain to a transformed economy is making the jump the data has not yet made.
What nobody has measured#
Output quality alongside speed, in the same study, in real work. Peng declined to measure it. The government trials could not. Noy and Zhang measured it and found it rose, on one task type, graded by evaluators.
Where the saved time went. GDS looked and could not tell. That question is the subject of the unclaimed hour, and nowhere has an answer.
And anything longer than a few months. Every study here is a snapshot. Whether a team that is faster this quarter is still faster in three years, or has quietly lost the understanding that made it fast, is the question this estate exists to ask and the one nothing yet answers.
What to take from this if you are deciding#
Expect large gains where the task is specified, bounded and new, and expect little or nothing where the expensive part is knowing the system. Measure output rather than asking about it, because the people doing the work cannot reliably tell you, in either direction. And if the case rests on minutes saved, decide in advance what those minutes are for, because the only organisation that has looked could not find out afterwards.
Key sources
- Noy, S. and Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence. Science. Graded entry.
- Peng, S., Kalliamvakou, E., Cihon, P. and Demirer, M. (2023). The Impact of AI on Developer Productivity: Evidence from GitHub Copilot. arXiv:2302.06590. Graded entry.
- Becker, J., Rush, N., Cunningham, T., Rein, D. and Mahamud, K. (2026). We are Changing our Developer Productivity Experiment Design. METR, 24 February 2026. Graded entry.
- Government Digital Service (2025). M365 Copilot Experiment: Cross-Government Findings Report. June 2025. Graded entry.
- Arzilli, F., Lynch, C. S. and Page, L. (2026). An Evaluation of DWP's Microsoft 365 Copilot Trial. DWP, 29 January 2026. Graded entry.
- Acemoglu, D. (2024). The Simple Macroeconomics of AI. Graded entry.
Related SuperSkills research#
The METR result in full is at what is the METR study, and the boundary problem at the jagged frontier. On measuring adoption without fooling yourself, how do you measure AI adoption properly and usage theatre. On the time itself, the unclaimed hour, and on who keeps the gain, who captures the productivity gains from AI.
About this research#
Written by Rahim Hirji, author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. Every figure on this page is kept with the study design that produced it, and self-reported figures are labelled as such throughout. Findings are attributed to the researchers who produced them and kept separate from the interpretation.
Evidence review · SS-2026-214 · Graded against the published rubric
Hirji, R. (2026). Does AI actually make people more productive?. The SuperSkills evidence base, SS-2026-214. https://thesuperskills.com/research/does-ai-actually-make-people-more-productive. Last reviewed 11 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work