The METR study is the randomised trial that found experienced developers took 19 per cent longer to finish real work when they were allowed to use AI, while believing the tools had made them faster. It has become the most quoted number in the argument about AI and productivity. In February 2026 METR put a warning banner on their own page saying the result is out of date, and most of the people quoting it have not noticed.
The answer, in one line
The METR study is a randomised controlled trial published in July 2025. Sixteen experienced open-source developers worked on 246 real issues from their own repositories, each issue randomly assigned to permit or forbid AI tools.
Definition#
The METR study: the randomised controlled trial published in July 2025 by METR, Model Evaluation and Threat Research, in which sixteen experienced open-source developers took 19 per cent longer to complete real tasks when permitted to use AI tools, having forecast a 24 per cent speed-up.
Sixteen developers, 246 issues, and a stopwatch#
METR recruited sixteen developers who had contributed for years to large open-source repositories, averaging more than 22,000 stars and a million lines of code. Each supplied a list of real issues from their own projects, 246 in total, of the kind they would have worked on anyway: bug fixes, features, refactors. Every issue was then randomly assigned to a condition permitting AI tools or forbidding them. Tasks averaged about two hours. Developers recorded their screens and reported the implementation time. They were paid 150 dollars an hour. The tooling was early 2025: mostly Cursor Pro running Claude 3.5 and 3.7 Sonnet, frontier models at the time.
The paper is arXiv:2507.09089, second version 25 July 2025, 51 pages. It carries no journal reference after more than a year, so it remains a preprint. Graded entry.
Forty points between what they felt and what the clock showed#
Before starting, the developers expected AI to speed them up by 24 per cent. Measured, they were 19 per cent slower. Afterwards, having lived through the slowdown themselves, they still estimated that AI had made them about 20 per cent faster. METR later put the size of the error plainly: participants overestimated AI's effect on their time by 40 percentage points on average.
That gap is the reason this study matters to anyone outside software. It is a controlled demonstration that a competent professional, working on their own material, can be wrong about the direction of their own productivity. Not wrong about the size of a gain. Wrong about whether there was one.
METR investigated twenty candidate explanations for the slowdown and found evidence that five contributed. They ruled out the obvious experimental artefacts: developers used frontier models, complied with their assignment, did not selectively drop hard tasks from the AI-disallowed arm, and submitted pull requests of similar quality either way.
METR have put a warning banner on their own result#
On 24 February 2026, Becker, Rush, Cunningham, Rein and Mahamud published an update. The original page now opens with a warning that the results are out of date and that METR believe they no longer reflect the current impact of AI models on open-source developer productivity. Graded entry.
The second experiment began in August 2025 with 57 developers, ten returning from the first study and 47 newly recruited across 143 repositories and more than 800 tasks, at 50 dollars an hour instead of 150. The raw numbers now point the other way. Returning developers show an estimated speed-up of 18 per cent, with a confidence interval running from a 38 per cent speed-up to a 9 per cent slowdown. New recruits show 4 per cent, interval from 15 per cent faster to 9 per cent slower. Every one of those intervals crosses zero, and METR say so.
Why the second trial could not settle it#
The reason METR give for distrusting their own new data is the most useful thing in the update. Between 30 and 50 per cent of developers told them they were declining to submit particular tasks, because they did not want to do those tasks without AI. An increased share declined to take part at all. One participant described avoiding issues where AI would finish in two hours what would otherwise take twenty.
So the experiment is systematically missing the tasks with the highest expected uplift and the people with the highest expectations. METR treat their estimate as a lower bound and are redesigning the study. The structural point generalises well beyond software: as a technology becomes normal, the population willing to be measured without it stops resembling the population using it. Every controlled trial of AI at work will meet this, and it gets harder each year rather than easier.
What METR measured instead, and the subgroup that reported least#
Between February and April 2026, Joel Becker surveyed 349 technical workers: 87 software engineers, 71 researchers, 129 academics and PhD students, 48 founders and managers. The survey deliberately asked about value produced rather than speed, on the argument that speed gains overstate value gains when AI changes which tasks a person takes on at all. Graded entry.
Median self-reported value change was between 1.4 and 2 times. Median self-reported speed change was 3 times, which is the gap the design predicted. Asked the same question about different years, respondents put themselves at 1.3 times in March 2025, 2 times in March 2026, and forecast 2.5 times for March 2027.
One result inside that survey deserves more attention than the headline. METR's own staff gave the lowest value-change answers of any subgroup studied, and METR suggest the reason is that those staff have the perception-gap finding in mind when they answer. Knowing about the measurement error appears to shrink the reported gain. That is a survey observation on a small subgroup and not a designed test, so it proves nothing on its own. Coming from the people who ran the original trial, it is still the most interesting sentence METR published this year.
The sample is a convenience sample drawn from GitHub, academic directories, METR and its staff's professional networks, with an email response rate around 2 per cent and about 70 per cent of participants paid, averaging 200 dollars. METR name the selection bias themselves.
Speed was the wrong quantity all along#
METR's move from measuring speed to measuring value is the interesting part of this sequence, and it arrives at the question this research has been asking from the other direction. In Box of Amazing on 29 March 2026, Rahim Hirji wrote that the brave question is not whether AI will take your job but "how much of what I call 'my job' is actually just friction I've learned to live with?", and that the better question about any new technology is "what would I do if this took no time at all?" rather than how to do the old thing faster.
An organisation counting hours saved is measuring the first phase. The METR sequence shows what happens to that measurement when you look closely: the self-report runs high, the controlled measurement is hard to run at all, and the quantity everybody reports is the one the researchers now think least worth having. This is the evidential floor under usage theatre and under the unclaimed hour, which asks where the saved time actually went.
Four claims the trial does not support#
METR published these themselves, in a table, on the day of release. The study is not evidence that AI fails to speed up most software developers, since sixteen people on mature repositories represent nobody but themselves. It is not evidence about any domain other than software. It is not evidence about future systems. And it is not evidence that better use of the same tools could not have produced a speed-up in the same setting.
Two further limits belong to this page rather than to METR. The measured outcome is self-reported implementation time, screen-recorded but reported by the participant, so the instrument is not wholly independent of the person. And the study measures time, not the quality of what was produced or what the developer could still do a year later, which is the question deskilling asks.
If you are about to quote 19 per cent#
Date it. The figure belongs to early-2025 tooling in one setting, and METR have withdrawn it as a current signal. A slide that presents it as the state of AI productivity in 2026 is making a claim its own source disowns.
Quote the perception gap instead, because that is the part nothing has overturned. The second study did not touch it, the 2026 survey reproduces its logic. For a board the finding reads like this: the people using the tools cannot tell you what the tools are doing to their output, and asking them harder will not fix it. If a productivity case rests on self-reported time savings, it rests on the one quantity this literature has shown to be unreliable by 40 percentage points.
Then measure something that survives selection. Counting what a team can still do without the system is harder to game than counting hours saved. That counting is the design behind a capability audit.
Key sources
- Becker, J., Rush, N., Barnes, E. and Rein, D. (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. METR, 10 July 2025; paper at arXiv:2507.09089. Graded entry.
- Becker, J., Rush, N., Cunningham, T., Rein, D. and Mahamud, K. (2026). We are Changing our Developer Productivity Experiment Design. METR, 24 February 2026. Graded entry.
- Becker, J. (2026). Measuring the Self-Reported Impact of Early-2026 AI on Technical Worker Productivity. METR, 11 May 2026. Graded entry.
Related SuperSkills research#
On the gap between reported and real adoption, usage theatre and how to measure AI adoption properly. On where saved time goes, the unclaimed hour. On the uneven capability the trial ran into, the jagged frontier. On numbers that circulate past their evidence, the most quoted AI statistics, checked. On what a self-report cannot see, the illusion of competence.
About this research#
Written by Rahim Hirji, author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. METR is an independent research nonprofit and the study, the withdrawal and the survey are theirs. Nothing on this page is a SuperSkills coinage. Every figure here was read on METR's own pages and on arXiv, and the interpretation is kept separate from the findings.
Explainer · SS-2026-209 · Graded against the published rubric
Hirji, R. (2026). What is the METR study?. The SuperSkills evidence base, SS-2026-209. https://thesuperskills.com/research/what-is-the-metr-study. Last reviewed 9 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work