← Research
Research

What is the METR study?

The most quoted number in the argument about AI and productivity, and the organisation that produced it has marked it out of date.

Last reviewed: 9 September 2026

The trial design, the forty percentage points between what developers felt and what the clock showed, METR's own withdrawal of February 2026, the survey that replaced it, and the finding that outlives the number.

Questions this page answersAll 811 questions this research covers

The METR study is the randomised trial that found experienced developers took 19 per cent longer to finish real work when they were allowed to use AI, while believing the tools had made them faster. It has become the most quoted number in the argument about AI and productivity. In February 2026 METR put a warning banner on their own page saying the result is out of date, and most of the people quoting it have not noticed.

The answer, in one line

The METR study is a randomised controlled trial published in July 2025. Sixteen experienced open-source developers worked on 246 real issues from their own repositories, each issue randomly assigned to permit or forbid AI tools.

Share as a card

Definition#

The METR study: the randomised controlled trial published in July 2025 by METR, Model Evaluation and Threat Research, in which sixteen experienced open-source developers took 19 per cent longer to complete real tasks when permitted to use AI tools, having forecast a 24 per cent speed-up.

Share this definition as a card

Sixteen developers, 246 issues, and a stopwatch#

METR recruited sixteen developers who had contributed for years to large open-source repositories, averaging more than 22,000 stars and a million lines of code. Each supplied a list of real issues from their own projects, 246 in total, of the kind they would have worked on anyway: bug fixes, features, refactors. Every issue was then randomly assigned to a condition permitting AI tools or forbidding them. Tasks averaged about two hours. Developers recorded their screens and reported the implementation time. They were paid 150 dollars an hour. The tooling was early 2025: mostly Cursor Pro running Claude 3.5 and 3.7 Sonnet, frontier models at the time.

The paper is arXiv:2507.09089, second version 25 July 2025, 51 pages. It carries no journal reference after more than a year, so it remains a preprint. Graded entry.

Forty points between what they felt and what the clock showed#

Before starting, the developers expected AI to speed them up by 24 per cent. Measured, they were 19 per cent slower. Afterwards, having lived through the slowdown themselves, they still estimated that AI had made them about 20 per cent faster. METR later put the size of the error plainly: participants overestimated AI's effect on their time by 40 percentage points on average.

That gap is the reason this study matters to anyone outside software. It is a controlled demonstration that a competent professional, working on their own material, can be wrong about the direction of their own productivity. Not wrong about the size of a gain. Wrong about whether there was one.

METR investigated twenty candidate explanations for the slowdown and found evidence that five contributed. They ruled out the obvious experimental artefacts: developers used frontier models, complied with their assignment, did not selectively drop hard tasks from the AI-disallowed arm, and submitted pull requests of similar quality either way.

METR have put a warning banner on their own result#

On 24 February 2026, Becker, Rush, Cunningham, Rein and Mahamud published an update. The original page now opens with a warning that the results are out of date and that METR believe they no longer reflect the current impact of AI models on open-source developer productivity. Graded entry.

The second experiment began in August 2025 with 57 developers, ten returning from the first study and 47 newly recruited across 143 repositories and more than 800 tasks, at 50 dollars an hour instead of 150. The raw numbers now point the other way. Returning developers show an estimated speed-up of 18 per cent, with a confidence interval running from a 38 per cent speed-up to a 9 per cent slowdown. New recruits show 4 per cent, interval from 15 per cent faster to 9 per cent slower. Every one of those intervals crosses zero, and METR say so.

Why the second trial could not settle it#

The reason METR give for distrusting their own new data is the most useful thing in the update. Between 30 and 50 per cent of developers told them they were declining to submit particular tasks, because they did not want to do those tasks without AI. An increased share declined to take part at all. One participant described avoiding issues where AI would finish in two hours what would otherwise take twenty.

So the experiment is systematically missing the tasks with the highest expected uplift and the people with the highest expectations. METR treat their estimate as a lower bound and are redesigning the study. The structural point generalises well beyond software: as a technology becomes normal, the population willing to be measured without it stops resembling the population using it. Every controlled trial of AI at work will meet this, and it gets harder each year rather than easier.

What METR measured instead, and the subgroup that reported least#

Between February and April 2026, Joel Becker surveyed 349 technical workers: 87 software engineers, 71 researchers, 129 academics and PhD students, 48 founders and managers. The survey deliberately asked about value produced rather than speed, on the argument that speed gains overstate value gains when AI changes which tasks a person takes on at all. Graded entry.

Median self-reported value change was between 1.4 and 2 times. Median self-reported speed change was 3 times, which is the gap the design predicted. Asked the same question about different years, respondents put themselves at 1.3 times in March 2025, 2 times in March 2026, and forecast 2.5 times for March 2027.

One result inside that survey deserves more attention than the headline. METR's own staff gave the lowest value-change answers of any subgroup studied, and METR suggest the reason is that those staff have the perception-gap finding in mind when they answer. Knowing about the measurement error appears to shrink the reported gain. That is a survey observation on a small subgroup and not a designed test, so it proves nothing on its own. Coming from the people who ran the original trial, it is still the most interesting sentence METR published this year.

The sample is a convenience sample drawn from GitHub, academic directories, METR and its staff's professional networks, with an email response rate around 2 per cent and about 70 per cent of participants paid, averaging 200 dollars. METR name the selection bias themselves.

Speed was the wrong quantity all along#

METR's move from measuring speed to measuring value is the interesting part of this sequence, and it arrives at the question this research has been asking from the other direction. In Box of Amazing on 29 March 2026, Rahim Hirji wrote that the brave question is not whether AI will take your job but "how much of what I call 'my job' is actually just friction I've learned to live with?", and that the better question about any new technology is "what would I do if this took no time at all?" rather than how to do the old thing faster.

An organisation counting hours saved is measuring the first phase. The METR sequence shows what happens to that measurement when you look closely: the self-report runs high, the controlled measurement is hard to run at all, and the quantity everybody reports is the one the researchers now think least worth having. This is the evidential floor under usage theatre and under the unclaimed hour, which asks where the saved time actually went.

Four claims the trial does not support#

METR published these themselves, in a table, on the day of release. The study is not evidence that AI fails to speed up most software developers, since sixteen people on mature repositories represent nobody but themselves. It is not evidence about any domain other than software. It is not evidence about future systems. And it is not evidence that better use of the same tools could not have produced a speed-up in the same setting.

Two further limits belong to this page rather than to METR. The measured outcome is self-reported implementation time, screen-recorded but reported by the participant, so the instrument is not wholly independent of the person. And the study measures time, not the quality of what was produced or what the developer could still do a year later, which is the question deskilling asks.

If you are about to quote 19 per cent#

Date it. The figure belongs to early-2025 tooling in one setting, and METR have withdrawn it as a current signal. A slide that presents it as the state of AI productivity in 2026 is making a claim its own source disowns.

Quote the perception gap instead, because that is the part nothing has overturned. The second study did not touch it, the 2026 survey reproduces its logic. For a board the finding reads like this: the people using the tools cannot tell you what the tools are doing to their output, and asking them harder will not fix it. If a productivity case rests on self-reported time savings, it rests on the one quantity this literature has shown to be unreliable by 40 percentage points.

Then measure something that survives selection. Counting what a team can still do without the system is harder to game than counting hours saved. That counting is the design behind a capability audit.

Key sources

On the gap between reported and real adoption, usage theatre and how to measure AI adoption properly. On where saved time goes, the unclaimed hour. On the uneven capability the trial ran into, the jagged frontier. On numbers that circulate past their evidence, the most quoted AI statistics, checked. On what a self-report cannot see, the illusion of competence.

About this research#

Written by Rahim Hirji, author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. METR is an independent research nonprofit and the study, the withdrawal and the survey are theirs. Nothing on this page is a SuperSkills coinage. Every figure here was read on METR's own pages and on arXiv, and the interpretation is kept separate from the findings.

How this research works  ·  Reviewed quarterly  ·  Found an error? Tell me and it is corrected on the page.

Explainer · SS-2026-209 · Graded against the published rubric

Cite this page

Hirji, R. (2026). What is the METR study?. The SuperSkills evidence base, SS-2026-209. https://thesuperskills.com/research/what-is-the-metr-study. Last reviewed 9 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

What is the METR study?

The METR study is a randomised controlled trial published in July 2025. Sixteen experienced open-source developers worked on 246 real issues from their own repositories, each issue randomly assigned to permit or forbid AI tools. Developers took 19 per cent longer when AI was permitted, having forecast a 24 per cent speed-up, and afterwards still believed AI had made them about 20 per cent faster. The paper is arXiv:2507.09089 and remains a preprint.

What does METR stand for?

Model Evaluation and Threat Research, which METR gives as its own subtitle on its about page. It is pronounced meter. METR is an independent research nonprofit that evaluates frontier AI models, and the developer productivity work is one strand of a wider programme on AI capabilities and risks.

Is the 19 per cent slowdown still true?

METR say it is out of date. On 24 February 2026 they published new data and added a warning to the original page saying the historical results no longer reflect the current impact of AI on open-source developer productivity. Their second experiment gives raw estimates pointing the other way, an 18 per cent speed-up for returning developers and 4 per cent for new recruits, but every confidence interval crosses zero and METR treat the whole estimate as an unreliable signal because 30 to 50 per cent of developers declined to submit tasks they did not want to do without AI.

Did METR retract the study?

No. They withdrew the 19 per cent as a statement about the present and left the study standing as a measurement of early-2025 tooling. The perception-gap finding is untouched: participants overestimated AI's effect on their time by 40 percentage points on average, a figure METR restated in May 2026. Quote the perception gap rather than the slowdown.

In this hub

Definitions

The terms this field uses, defined against their primary sources.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory for CEOs and boards  ·  Enquire