← Research
Research · Definition

What is the METR time horizon?

A benchmark statistic with its success rate written into its name, quoted almost everywhere without it.

Last reviewed: 22 September 2026 · Next review due: 22 September 2027

The metric Kwa and colleagues proposed in March 2025, how the 207-day doubling time was built, the 80 per cent horizon that runs four to six times shorter, the messiness score the authors published against themselves, and why this is a different METR result from the developer trial.

Questions this page answersAll 887 questions this research covers

Definition#

The METR time horizon is the length of task, measured by how long a human expert takes to do it, that an AI model completes about half the time. Kwa and colleagues proposed the metric in March 2025 and named it the 50%-task-completion time horizon. On their own task suite the frontier models of that moment sat at around fifty minutes.

Share this definition as a card

The success rate sits inside the name of the metric and falls out of almost every quotation of it. A fifty-minute time horizon describes the tasks a model finishes on a coin flip. The same authors measured the horizon at 80 per cent success and found it four to six times shorter, and they say in the paper that they cannot measure it at all at the reliability levels an organisation would actually deploy at. Ask anyone quoting a horizon at what success rate, and about half the argument goes away.

The answer, in one line

The METR time horizon is the length of task that an AI model completes about half the time, where task length is measured by how long a human expert takes to do the same task.

Share as a card

The metric, in the authors' own words#

METR, Model Evaluation and Threat Research, proposed the measure to solve a problem it names in its first line: "Despite rapid progress on AI benchmarks, the real-world meaning of benchmark performance remains unclear." A score of 64 per cent on a benchmark tells a reader nothing about what a system can be handed. Time does, because everyone has an intuition for it.

So they anchored capability to human clock time: "the duration of tasks that models can complete at a certain success probability, providing an intuitive measure of real-world capability compared to humans. As models may not reliably complete all tasks of a given length, we operationalize this by measuring the X%-(task completion) time horizon, the length of tasks that models can complete approximately X% of the time."

The construction is a logistic regression per model. Each task is given a length equal to the geometric mean time taken by the humans who succeeded at it, and each model is fitted a horizon at which its probability of success crosses one half. GPT-2 comes out at 2 seconds. o3 comes out at 110 minutes and "succeeds at several tasks over 4 hours". In between, from their own table, Claude 3 Opus is 6.42 minutes, GPT-4o 9.17, the November 2024 Claude 3.5 Sonnet 28.98 and o1 39.21.

Where the minutes come from is more interesting than the minutes#

The denominator of this metric is human beings with a stopwatch on them, and the suite they were timed on is narrow by design. It holds 170 tasks in three groups: 97 software tasks from HCAST running between one minute and thirty hours, seven machine-learning research engineering tasks from RE-Bench all eight hours long, and 66 software atomic actions of one to thirty seconds written to cover the short end. The short ones exist because a metric measured in minutes needs something below a minute to sit on.

The human side runs to "over 800 baselines totaling 2,529 hours", of which 286 successful ones come from the longer tasks and 236 from the short ones. Twenty-one of the 169 tasks have no human baseline at all and carry a researcher's estimate instead. The baseliners are described as skilled professionals in software engineering, machine learning and cybersecurity, mostly from top-100 universities, with about five years of relevant experience. They were recorded, and paid bonuses for finishing quickly.

Two admissions in the appendices change how the headline should be read. Conditioning on success shortens the measured task length, "thereby underestimating model performance", worst on the long tasks where most baseliners fail. And when METR compared their contract baseliners against the people who maintain the repositories in question, the contractors took five to eighteen times longer. Their own conclusion: "time horizons may have better correspondence to the labor of a low-context human, rather than a high-context human". A fifty-minute horizon is fifty minutes of somebody arriving cold.

Seven months, with a fifth of slack either side#

The finding that travels is the doubling time. Across 2019 to 2025 the horizon doubled "every 207 days with a 95% bootstrapped confidence interval 166-240 days (roughly ±19%)", from a three-level hierarchical bootstrap over task families, tasks and runs. Seven months is 207 days rounded, and the rounding is where the uncertainty went.

The 80 per cent horizon doubles at almost the same rate, 204 days, which is the part of the result that holds up best: the slope survives a change in the threshold even though the level does not. The authors put it plainly themselves. "While there are wide error bars on each individual models' horizon lengths, these errors are highly correlated between models... Therefore, we are more confident in the slope of the time horizon trend than in the time horizon of any particular model." Anyone quoting a single model's horizon is quoting the weaker half of the paper.

The acceleration claim needs more care again. Fitting only models released in 2024 produces a doubling time of about three months, and the authors immediately add that "any extrapolation into the future would not be robust". On the 2024 to 2025 window they write: "Note our confidence in the 2024-2025 only trend is low because there are only seven frontier models in this time span." They defer the significance test for a slope change to future work twice, in two separate appendices, which is a way of saying they have not run it.

The authors scored their own tasks for messiness and published the number#

The obvious objection to any benchmark is that real work is not like this. METR raised it against themselves and put a figure on it. They labelled every task for the presence or absence of sixteen "messiness" factors, the things that make real work hard: being under-specified, unclear feedback loops, coordination between streams of work in real time. The mean score across their suite is 3.2 out of 16, none of the tasks exceeds 8, and for comparison they estimate that "a task like 'write a good research paper' would score between 9/16 and 15/16".

Messiness costs models performance. Controlling for length, "an increase in task messiness by 1 point reduces mean success rates by roughly 8.1%". So a paper-writing task at 12 out of 16 sits nine points above the suite mean, on a gradient that the paper itself fitted.

The result runs the other way too, and the page would be dishonest to omit it. Splitting the suite into messier and cleaner halves, the rate of improvement over time is the same in both: "success rates increased by 40 percentage points between Jan. 2023 and May 2025 in both high and low messiness splits", and they "find no evidence of either much slower performance trends, or a plateau, specific to our higher messiness subset". Messiness moves the level and has not so far moved the slope. That is a real finding against the comfortable view, measured on tasks that top out at half the messiness of writing a paper.

Appendix B is where the external validity question is answered most directly, in a sentence the authors volunteer: "The tasks we use to benchmark AI capabilities are systematically different from real tasks. These differences could result in the trends we observe on these tasks not generalizing to real world tasks." Their five reasons are that everything is automatically scored, nothing involves other agents, almost nothing constrains a scarce resource, almost nothing punishes a single mistake, and the environments stay still unless the agent acts on them. Most professional work fails all five conditions before lunch.

Two METR results circulate under one name#

This estate has already corrected twelve pages for carrying a METR figure past its expiry. The confusion is easy to make and worth heading off, because the two results measure opposite ends of the same question and one of them has a warning attached.

The developer trial, arXiv:2507.09089, randomised 246 real issues across sixteen experienced open-source developers and found them 19 per cent slower with AI permitted while they believed themselves faster. METR marked that figure out of date on 24 February 2026. It is a measurement of what happened to people.

The time horizon, arXiv:2503.14499, is a measurement of what a model does on a benchmark with no humans in the loop at all. It carries no withdrawal and it has a different weakness: the number is quoted stripped of the success rate that defines it. One result was retired and keeps being cited; the other is live and keeps being cited wrong. Neither licenses a claim about the other.

The 80 per cent figure everyone reaches for does not exist#

The paper reports the 80 per cent horizon as a ratio and a doubling time, never as an absolute. The text gives "roughly 5x shorter" in the introduction and "4-6x shorter" in the body, plus the 204-day doubling. The minute values live only inside a chart. No model's 80 per cent horizon is stated anywhere in the body or the appendices.

Applying the ratio to o3's 110 minutes would land somewhere between eighteen and twenty-eight minutes, and that arithmetic is not the authors' and should not be attributed to them. It is written out here so that nobody has to reconstruct it privately. Neither this page nor any other on this site uses the result. The authors were equally direct about the ceiling: "Due to the limited dataset, we cannot confidently measure time horizons at very high success rates (e.g. 95%)". Measuring reliability properly, they note, "requires very large task datasets with near-zero label noise".

The paper changed its own title, and the old one is what circulates#

Version one, posted 18 March 2025, is called Measuring AI Ability to Complete Long Tasks and carries 25 authors. The current version four, posted 10 July 2026, is called Measuring AI Ability to Complete Long Software Tasks, carries 26 with Chris Painter added, and now gives a journal reference of NeurIPS 2025. Both abstract pages were read at source on 22 September 2026.

Three things follow. The paper is no longer a preprint and should not be graded as one. The authors narrowed their own claim to software, and the unnarrowed title is the one still attached to the figure in most secondary coverage. And their own footnote had said as much from the start: "our forecasts concern AI with a 1-month horizon on software tasks, not 1-month AGI, because we evaluate models only on software and research tasks." The title caught up with the footnote fifteen months later.

Reading a horizon without being sold one#

Three questions make the number usable, and all three are answerable from the paper rather than from whoever is quoting it.

At what success rate. If the answer is 50 per cent, the figure describes tasks the system gets right half the time, which for most work is a description of how much checking is required and not of what can be handed over. The estate's standing interest in the cost of verification starts exactly here: a system that is right half the time on hour-long tasks generates hour-long tasks to check.

On which tasks, and how messy. The suite is software and research engineering, automatically scored, with a messiness mean of 3.2 out of 16. A capability claim drawn from it applies to work of that shape. The authors expect horizons "to differ by a large factor depending on the task domain and reference human population", and say so twice.

Slope or level. The doubling time is the part of this paper its own authors trust most, and the individual model horizons are the part they trust least. A board briefing built on "frontier models can now do fifty-minute tasks" has taken the weak half. One built on "the length of task these systems can attempt has been doubling on a clock measured in months, on a benchmark that does not look much like our work" has taken the strong half and its caveat together.

For the wider question of how an organisation should treat any single capability figure, see how to measure AI adoption properly and why human in the loop is not a safeguard, which is the same observation arriving from the deployment side.

Key sources

On the other METR result and its withdrawal, what the METR study is. On the uneven shape of what these systems can do, the jagged frontier. On what checking costs once the output is long, the verifier's discount. On how the work actually divides, how humans and agents divide work. On the oversight this leaves behind, meaningful human oversight.

About this research#

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. The time horizon metric, the messiness score and every figure on this page belong to Kwa and colleagues at METR. Nothing here is a SuperSkills coinage. The paper was read at source in its version four HTML and both the version one and version four abstract pages on 22 September 2026. No absolute 80 per cent horizon is quoted, because the paper prints none. Last reviewed: 22 September 2026.

How this research works  ·  Reviewed quarterly  ·  Found an error? Tell me and it is corrected on the page.

Explainer · SS-2026-287 · Graded against the published rubric

Cite this page

Hirji, R. (2026). What is the METR time horizon?. The SuperSkills evidence base, SS-2026-287. https://thesuperskills.com/research/what-is-the-metr-time-horizon. Last reviewed 22 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

What is the METR time horizon?

The METR time horizon is the length of task that an AI model completes about half the time, where task length is measured by how long a human expert takes to do the same task. Kwa and colleagues proposed it in March 2025 and called it the 50%-task-completion time horizon. On their own suite, frontier models of that moment sat at around 50 minutes, GPT-2 at about 2 seconds, and o3 at 110 minutes.

Why does the 50 per cent matter?

Because the success rate is part of the definition and is usually dropped when the figure is quoted. A 50-minute horizon describes tasks the model finishes on a coin flip. The same authors measured the 80 per cent horizon and report that it runs four to six times shorter, writing that even models that sometimes succeed on difficult and diverse tasks cannot reliably perform tasks of moderate length. They say they cannot confidently measure horizons at very high success rates such as 95 per cent, because the dataset is too small.

How fast is the time horizon doubling?

Every 207 days across 2019 to 2025, with a 95 per cent bootstrapped confidence interval of 166 to 240 days. The 80 per cent horizon doubles on a similar clock, at 204 days. Restricting the fit to models released in 2024 gives roughly three months instead, and the authors add that on that basis "any extrapolation into the future would not be robust". They also state that their confidence in the 2024 to 2025 trend alone is low, because only seven frontier models fall inside it.

Is the time horizon the same as METR's 19 per cent slowdown study?

No. They are two different pieces of work from the same organisation. The time horizon is a benchmark metric for model capability, from arXiv:2503.14499. The 19 per cent figure comes from a randomised trial of sixteen developers working on their own repositories, arXiv:2507.09089, and METR marked that figure out of date in February 2026. Neither number tells you anything about the other, and the time horizon carries no withdrawal.

How representative are the tasks the horizon is measured on?

The authors answer this against themselves. They write that the tasks are systematically different from real tasks, and list why: everything is automatically scored, nothing involves other agents, few tasks constrain a scarce resource, few punish a single mistake, and the environments do not change unless the agent acts on them. They also scored their own tasks on sixteen messiness factors and report a mean of 3.2 out of 16, with none above 8, against an estimate of 9 to 15 for a task like writing a good research paper.

In this hub

Definitions

The terms this field uses, defined against their primary sources.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory for CEOs and boards  ·  Enquire