← Research
Research

Does AI actually make people more productive?

Large gains on narrow tasks, scattered results in real work, and no sign of it yet in the economy.

Last reviewed: 11 September 2026

Six measurements that found large gains, two that found none or worse, two government trials of 23,500 licences that could only ask people, and the perception gap that runs in both directions.

Questions this page answersAll 811 questions this research covers

Yes, on narrow tasks, by amounts large enough that arguing about them is pointless. Inside real work the answer scatters, and one careful randomised study measured experienced people going slower. At the level of an economy, nothing has shown up yet. This page keeps those three altitudes apart, because almost every argument about AI and productivity is two people describing different ones.

The answer, in one line

On narrow, well-specified tasks the measured gains are large and repeated: writing tasks completed 40 per cent faster with 18 per cent higher rated quality, a standardised coding task 55.8 per cent faster, customer support resolutions up 15 per cent an hour.

Share as a card

The short answer#

On well-specified tasks with a clear finish line, generative AI produces large measured gains. In continuing work inside systems people already know, the measured effect is uncertain and sometimes negative. Nothing has yet appeared in national productivity statistics.

Share this definition as a card

Where the gains are large and well measured#

Noy and Zhang randomised professionals across mid-level writing tasks. Average time fell 40 per cent and rated output quality rose 18 per cent, with the gap between stronger and weaker writers narrowing. Graded entry.

Peng and colleagues ran a randomised trial of GitHub Copilot on 95 professional programmers implementing an HTTP server in JavaScript. Among those who finished, the assisted group took 71 minutes against 161, a 55.8 per cent reduction, p=0.0017. The authors state plainly that they did not examine code quality. Graded entry.

Brynjolfsson, Li and Raymond studied 5,000 customer support agents in the field. Resolutions per hour rose about 15 per cent overall, 30 per cent for the least experienced and 36 per cent for the lowest skill quintile, while the most skilled showed no significant gain. Graded entry.

Three different task types, three large effects, one consistent pattern underneath: the gains land hardest on the people who were furthest from the standard, which is the compression result that runs through this whole literature.

Where they thin out, and where they reverse#

Dell'Acqua and colleagues gave 758 consultants GPT-4 across two task types. Inside the model's competence the assisted consultants were dramatically better. On a task designed to sit just outside it, they did worse than consultants with no AI at all. The authors call the boundary a jagged technological frontier, and its importance is that nobody can see where it runs from inside the work. Graded entry.

METR's 2025 experiment is the one that broke the pattern. Experienced open-source developers, working on their own mature repositories, were measured 19 per cent slower when permitted to use AI tools. The tasks were real, the repositories were ones they knew well, and the finding survives as the most cited counter-result in the field. Graded entry.

Set the two coding studies beside each other and the shape of the disagreement appears. Peng's task was a fresh server built from nothing against a clear specification. METR's tasks were changes to large codebases their authors had lived in for years. Speed arrives where the work is production. It disappears where the work is understanding.

METR has since unsettled its own number, and said so#

In February 2026 METR reported a second randomised study, 57 developers across 143 repositories and more than 800 tasks, and stated that their data now gives an unreliable signal. Between 30 and 50 per cent of developers declined to submit tasks they did not want to do without AI, which selects the sample. The raw results point the other way, a speedup of -18 per cent for returning developers and -4 per cent for new recruits, and every confidence interval crosses zero. They believe developers are probably more sped up in 2026 than in 2025 and say their own data is only very weak evidence for it. Graded entry.

Two things follow, and both matter more than the number. The original study is not retracted, and its perception finding is untouched. Anyone still quoting "19 per cent slower" as the current state of AI coding tools is quoting early-2025 tooling and a design its own authors have replaced.

The perception gap runs in both directions#

The most quoted line from METR is that developers forecast a 24 per cent speed-up, were measured 19 per cent slower, and after finishing the tasks still estimated AI had made them about 20 per cent faster. It is a striking result and it has hardened into a general law: people overstate what AI does for them.

The law does not hold. In the Copilot trial, participants in both arms estimated a 35 per cent improvement when the measured figure was 55.8 per cent, so they understated it by twenty points. METR's own 2026 survey team, reporting a median self-reported speed change of three times, note that only one study has gathered survey and field-experiment results on the same population and metric, and decline to say how far surveys overstate in general. Graded entry.

So the honest statement is narrower and more useful than the slogan. People are poor estimators of their own throughput, in whichever direction the error happens to fall, and asking them is not a measurement. That is a reason to measure, and not a reason to assume the answer is smaller than they say.

Two government trials, 23,500 licences, and no baseline between them#

The Government Digital Service ran 20,000 Copilot licences across twelve organisations for three months. Average self-reported saving was 26 minutes a day, adoption held around 80 per cent, and 82 per cent said they would not want to go back. The report's own conclusions state that it was not possible to identify how the saved time was spent. Graded entry.

The Department for Work and Pensions evaluated 3,549 licences against a stratified comparison group of non-users, and estimated 19 minutes a day across eight routine tasks, statistically significant across specifications, with job satisfaction up 0.56 points. Its own limitations chapter names the absence of baseline data, post-treatment bias and self-selection towards enthusiasts through first-come first-served allocation, which it says may lead to overestimation. Graded entry.

These are the largest organisational deployments anybody has published on, and neither can tell you whether the organisation produced more. They are careful, useful documents about how people experienced a tool. An estimate of 26 minutes, multiplied by a headcount, is a business case built on a survey question, and both departments are more honest about that than the people quoting them.

The economy has not noticed#

Acemoglu's estimate remains the most careful macroeconomic one available: total factor productivity gains of no more than 0.66 per cent over ten years, revised to under 0.53 per cent once the difficulty of hard-to-learn tasks is accounted for. Graded entry.

A gap this wide between task-level results and aggregate ones is not itself surprising, and it has a long history in economics. It does mean that anyone extrapolating from a 40 per cent writing gain to a transformed economy is making the jump the data has not yet made.

What nobody has measured#

Output quality alongside speed, in the same study, in real work. Peng declined to measure it. The government trials could not. Noy and Zhang measured it and found it rose, on one task type, graded by evaluators.

Where the saved time went. GDS looked and could not tell. That question is the subject of the unclaimed hour, and nowhere has an answer.

And anything longer than a few months. Every study here is a snapshot. Whether a team that is faster this quarter is still faster in three years, or has quietly lost the understanding that made it fast, is the question this estate exists to ask and the one nothing yet answers.

What to take from this if you are deciding#

Expect large gains where the task is specified, bounded and new, and expect little or nothing where the expensive part is knowing the system. Measure output rather than asking about it, because the people doing the work cannot reliably tell you, in either direction. And if the case rests on minutes saved, decide in advance what those minutes are for, because the only organisation that has looked could not find out afterwards.

Key sources

The METR result in full is at what is the METR study, and the boundary problem at the jagged frontier. On measuring adoption without fooling yourself, how do you measure AI adoption properly and usage theatre. On the time itself, the unclaimed hour, and on who keeps the gain, who captures the productivity gains from AI.

About this research#

Written by Rahim Hirji, author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. Every figure on this page is kept with the study design that produced it, and self-reported figures are labelled as such throughout. Findings are attributed to the researchers who produced them and kept separate from the interpretation.

How this research works  ·  Reviewed quarterly  ·  Found an error? Tell me and it is corrected on the page.

Evidence review · SS-2026-214 · Graded against the published rubric

Cite this page

Hirji, R. (2026). Does AI actually make people more productive?. The SuperSkills evidence base, SS-2026-214. https://thesuperskills.com/research/does-ai-actually-make-people-more-productive. Last reviewed 11 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

Does AI actually make people more productive?

On narrow, well-specified tasks the measured gains are large and repeated: writing tasks completed 40 per cent faster with 18 per cent higher rated quality, a standardised coding task 55.8 per cent faster, customer support resolutions up 15 per cent an hour. Inside real work the picture scatters, and one randomised study measured experienced developers as 19 per cent slower. At the level of the economy nothing has yet appeared: the most careful estimate puts total factor productivity gains at no more than 0.66 per cent over ten years.

Do people overestimate how much AI speeds them up?

Sometimes, and the reverse also happens, which is rarely reported. In METR's 2025 experiment developers forecast a 24 per cent speed-up, were measured 19 per cent slower, and still believed afterwards that AI had made them about 20 per cent faster. In the GitHub Copilot trial participants estimated 35 per cent when the measured gain was 55.8 per cent, so they understated it. METR's own survey team decline the general claim that surveys always overstate.

Why do the studies disagree so much?

Because they measure different work. The large gains come from standardised tasks with a clear finish line, often started from nothing. The null and negative results come from experienced people working inside systems they already know well, where the expensive part is understanding rather than producing. Dell'Acqua's jagged frontier makes the same point directly: inside the model's competence AI-assisted consultants were dramatically better, and on a task just outside it they did worse than consultants with no AI at all.

What do the big organisational trials show?

They show what people say, not what changed. The UK Government Digital Service trialled 20,000 Copilot licences and found an average self-reported saving of 26 minutes a day, and the Department for Work and Pensions found 19 minutes a day across 3,549 licences with a non-user comparison group. Neither had baseline task timing, a randomised control or any measure of output quality, and the GDS report states it was not possible to identify how the saved time was spent.

In this hub

Work, careers and the labour market

What happens to jobs, careers and the first rung.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory for CEOs and boards  ·  Enquire