← Research
Research · Question

Will AI get better at this soon?

Measured capability is rising fast on scoreable tasks. Your task may not be one of them.

Last reviewed: 3 October 2026 · Next review due: 3 October 2027

What METR's time-horizon measurements, their stated caveats and two studies of reliability say about whether AI will get better at a given task, and how to test your own work each quarter.

Question this page answersAll 1156 questions this research covers

On the one family of tasks where progress is measured against human working time, yes, quickly. METR's January 2026 update puts the length of software task that frontier models complete half the time at roughly five hours for the best model it tested, and the trend has been shortening its doubling time. Whether your task improves depends on three things that measurement leaves out: whether it resembles a scored task, how reliably it has to be done, and whether the model was ever the limit. The authors who built the measure say they are more confident in the slope than in any single model's level.

The answer, in one line

On software tasks that can be scored automatically, measured capability is rising quickly: METR's January 2026 update gives a doubling time of 196 days for 2019 to 2025 and 89 days since 2024, with wide confidence intervals.

Share as a card

Definition#

Capability slope and capability level: the slope is how fast a measured capability rises over time, and the level is where one particular model sits on that measure at one moment. Kwa and colleagues at METR write that they are "more confident in the slope of the time horizon trend than in the time horizon of any particular model". The terms are descriptive and are not a SuperSkills coinage.

Share this definition as a card

The measured slope is steep, and METR's own estimates keep moving#

METR measures the human-expert task length that a model completes with about 50 per cent success, called the time horizon. The original paper, Kwa and colleagues, found it doubled every 207 days from 2019 to 2025. On 29 January 2026 METR released Time Horizon 1.1, which expanded the suite from 170 to 228 tasks and the number of tasks of eight hours or more from 14 to 31. Its estimates: for the full period, "exactly the same doubling time as the TH1 trend, of 196 days (7 months)"; since 2023, 131 days against 165 under the first version; since 2024, 89 days against 109. For the models it measured on the new suite, the 50 per cent horizon was 320 minutes for Claude Opus 4.5 (confidence interval 170 to 729), 214 for GPT-5 (117 to 480) and 121 for o3 (74 to 201).

METR flags the limits in the same post: "The trend in time horizon is somewhat sensitive to task composition." The "confidence intervals are still very wide", and it measured human baseline times for only 5 of its 31 long tasks. So the steepening since 2024 is a fit to a handful of models, and the first version's authors already called their 2024 to 2025 acceleration low-confidence for the same reason.

The measured tasks are not the tasks most people do#

Every task in the suite is scored automatically, and the authors state that the tasks are systematically different from real work: none involves other agents, few punish a single mistake, and the environments are static. In a July 2025 follow-up covering benchmarks in other domains, Kwa and Cheng found broadly similar rates of improvement to the seven-month doubling, and wrote that "the two most common highly-paid occupations in the US are management and nursing, the key skills of which are mostly unrepresented in these benchmarks." They also expect "all benchmarks to be easier than real-world tasks in the corresponding domain", since benchmark tasks are cheap to run and score. A steep slope on tasks that can be scored by a script says little about work where the standard is a judgement.

Half the time is the wrong reliability for most work#

The 50 per cent figure describes tasks completed on a coin flip. In the original paper the 80 per cent horizon was four to six times shorter in level while doubling on a similar clock, and the authors say they cannot confidently measure horizons at very high success rates such as 95 per cent. A task that must be right almost every time sits far below the headline minutes. Kalai, Nachum, Vempala and Zhang argue in an unrefereed 2025 preprint that models guess confidently because most benchmarks score an abstention no better than a wrong answer, which is a reason to expect stronger models to stay unreliable in a specific way: capability can rise while the habit of answering regardless survives.

A better model did not help the people in the medical trial#

In Bean and colleagues' preregistered trial, 1,298 UK adults used large language models for ten medical scenarios. The models alone identified the condition in 94.9 per cent of cases and participants using them identified a relevant condition in under 34.5 per cent. The limit was the exchange between person and model, and a stronger model would not obviously change it. The authors call their figures a lower bound for newer models, so the result does not say that improvement is irrelevant. It says the model's score was never the whole of the outcome. The case is set out on should I use AI for medical questions.

What the forecasts cannot tell you#

Waiting for a better model is a decision with a cost#

This section is interpretation, kept apart from the evidence above.

The question "will it get better soon" is often a way of postponing a decision about how to work. The evidence here supports two modest claims. Measured capability on scoreable tasks has been rising fast, so a task that fails today deserves a re-test every few months. And the part of the outcome that sits outside the model, the question asked, the checking, the standard applied, did not improve by itself in the one trial that tested it. That second claim is an inference from a single study. If it holds, people who have built the habit of testing a tool against their own work gain the most from each new release, because they can tell what changed. The practice is described on how do expert AI users work differently and the checking skill on how do I know when AI is wrong.

A quarterly re-test beats a forecast#

Key sources

On the metric, what is the METR time horizon. On error rates, how often is AI wrong. On the mismatch between benchmarks and work, what is the jagged frontier. On when a model improves a decision, when does AI improve a decision.

About this research#

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. The two METR posts were read at source on 3 October 2026, and the other three studies were read at source on 30 September and 1 October 2026; all are graded in the evidence base. Capability slope and capability level are descriptive terms and not SuperSkills coinages.

How this research works  ·  Reviewed quarterly  ·  Found an error? Tell me and it is corrected on the page.

Evidence review · SS-2026-388 · Graded against the published rubric · 2 working papers

Cite this page

Hirji, R. (2026). Will AI get better at this soon?. The SuperSkills evidence base, SS-2026-388. https://thesuperskills.com/research/will-ai-get-better-at-this-soon. Last reviewed 3 October 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

Will AI get better at this soon?

On software tasks that can be scored automatically, measured capability is rising quickly: METR's January 2026 update gives a doubling time of 196 days for 2019 to 2025 and 89 days since 2024, with wide confidence intervals. Whether your own task improves depends on whether it resembles those tasks and on how reliably it must be done.

How fast is AI improving?

METR's Time Horizon 1.1 reports the length of task models complete half the time doubling every 196 days over the full period, 131 days since 2023 and 89 days since 2024, and notes the trend is somewhat sensitive to task composition and the intervals are very wide.

Does a better AI model mean better results for me?

Not automatically. In the Bean trial, models alone identified the medical condition in 94.9 per cent of cases while people using them identified a relevant condition in under 34.5 per cent. The exchange between person and model was the limit.

What does a 50 per cent time horizon mean?

It is the length of task, measured by how long a human expert takes, that a model completes about half the time. At 80 per cent success the horizon was four to six times shorter in the original METR paper.

Should I wait for a better model before starting?

The evidence does not support waiting as a strategy. Re-testing a fixed set of your own tasks each quarter shows what has changed, and building the habit of checking results improves every later release's value. That second point is an inference from one trial.

In this hub

Definitions

The terms this field uses, defined against their primary sources.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory  ·  Boards  ·  Enquire