On the one family of tasks where progress is measured against human working time, yes, quickly. METR's January 2026 update puts the length of software task that frontier models complete half the time at roughly five hours for the best model it tested, and the trend has been shortening its doubling time. Whether your task improves depends on three things that measurement leaves out: whether it resembles a scored task, how reliably it has to be done, and whether the model was ever the limit. The authors who built the measure say they are more confident in the slope than in any single model's level.
The answer, in one line
On software tasks that can be scored automatically, measured capability is rising quickly: METR's January 2026 update gives a doubling time of 196 days for 2019 to 2025 and 89 days since 2024, with wide confidence intervals.
Definition#
Capability slope and capability level: the slope is how fast a measured capability rises over time, and the level is where one particular model sits on that measure at one moment. Kwa and colleagues at METR write that they are "more confident in the slope of the time horizon trend than in the time horizon of any particular model". The terms are descriptive and are not a SuperSkills coinage.
The measured slope is steep, and METR's own estimates keep moving#
METR measures the human-expert task length that a model completes with about 50 per cent success, called the time horizon. The original paper, Kwa and colleagues, found it doubled every 207 days from 2019 to 2025. On 29 January 2026 METR released Time Horizon 1.1, which expanded the suite from 170 to 228 tasks and the number of tasks of eight hours or more from 14 to 31. Its estimates: for the full period, "exactly the same doubling time as the TH1 trend, of 196 days (7 months)"; since 2023, 131 days against 165 under the first version; since 2024, 89 days against 109. For the models it measured on the new suite, the 50 per cent horizon was 320 minutes for Claude Opus 4.5 (confidence interval 170 to 729), 214 for GPT-5 (117 to 480) and 121 for o3 (74 to 201).
METR flags the limits in the same post: "The trend in time horizon is somewhat sensitive to task composition." The "confidence intervals are still very wide", and it measured human baseline times for only 5 of its 31 long tasks. So the steepening since 2024 is a fit to a handful of models, and the first version's authors already called their 2024 to 2025 acceleration low-confidence for the same reason.
The measured tasks are not the tasks most people do#
Every task in the suite is scored automatically, and the authors state that the tasks are systematically different from real work: none involves other agents, few punish a single mistake, and the environments are static. In a July 2025 follow-up covering benchmarks in other domains, Kwa and Cheng found broadly similar rates of improvement to the seven-month doubling, and wrote that "the two most common highly-paid occupations in the US are management and nursing, the key skills of which are mostly unrepresented in these benchmarks." They also expect "all benchmarks to be easier than real-world tasks in the corresponding domain", since benchmark tasks are cheap to run and score. A steep slope on tasks that can be scored by a script says little about work where the standard is a judgement.
Half the time is the wrong reliability for most work#
The 50 per cent figure describes tasks completed on a coin flip. In the original paper the 80 per cent horizon was four to six times shorter in level while doubling on a similar clock, and the authors say they cannot confidently measure horizons at very high success rates such as 95 per cent. A task that must be right almost every time sits far below the headline minutes. Kalai, Nachum, Vempala and Zhang argue in an unrefereed 2025 preprint that models guess confidently because most benchmarks score an abstention no better than a wrong answer, which is a reason to expect stronger models to stay unreliable in a specific way: capability can rise while the habit of answering regardless survives.
A better model did not help the people in the medical trial#
In Bean and colleagues' preregistered trial, 1,298 UK adults used large language models for ten medical scenarios. The models alone identified the condition in 94.9 per cent of cases and participants using them identified a relevant condition in under 34.5 per cent. The limit was the exchange between person and model, and a stronger model would not obviously change it. The authors call their figures a lower bound for newer models, so the result does not say that improvement is irrelevant. It says the model's score was never the whole of the outcome. The case is set out on should I use AI for medical questions.
What the forecasts cannot tell you#
- Whether your task is on the curve. Software and some reasoning benchmarks are measured, and management, care and judgement-heavy work largely are not.
- How long. A doubling time fitted to a few models is a description of the past. METR's own estimates changed between versions and its authors' caution is in print.
- The reliability you need. No horizon above is stated at the success rate most organisations would deploy at.
- Anything beyond tasks. A task horizon says nothing about cost, accountability or whether people will use a tool well.
Waiting for a better model is a decision with a cost#
This section is interpretation, kept apart from the evidence above.
The question "will it get better soon" is often a way of postponing a decision about how to work. The evidence here supports two modest claims. Measured capability on scoreable tasks has been rising fast, so a task that fails today deserves a re-test every few months. And the part of the outcome that sits outside the model, the question asked, the checking, the standard applied, did not improve by itself in the one trial that tested it. That second claim is an inference from a single study. If it holds, people who have built the habit of testing a tool against their own work gain the most from each new release, because they can tell what changed. The practice is described on how do expert AI users work differently and the checking skill on how do I know when AI is wrong.
A quarterly re-test beats a forecast#
- Keep a fixed set of ten real tasks. Save the prompts and what a good result looks like, so you can run them on each new model and compare like with like.
- Score at the reliability you need. Count how many of ten pass without correction, since your threshold may be nearer 95 per cent than 50.
- Separate model failures from handling failures. When a result is poor, check whether the instruction, the supplied context or the model was the cause. Only the last improves with the next release.
- Ask a vendor which task and what success rate. A claim that a model can work for hours needs both stated, as the METR definition requires.
Key sources
- METR (2026). Time Horizon 1.1. METR blog, 29 January 2026. Graded entry.
- Kwa, T. et al. (2025). Measuring AI Ability to Complete Long Software Tasks. arXiv:2503.14499. Graded entry.
- Kwa, T. and Cheng, V. (2025). How Does Time Horizon Vary Across Domains? METR blog, 14 July 2025. Graded entry.
- Kalai, A. T., Nachum, O., Vempala, S. S. and Zhang, E. (2025). Why Language Models Hallucinate. arXiv:2509.04664, preprint. Graded entry.
- Bean, A. M. et al. (2026). Reliability of LLMs as medical assistants for the general public. Nature Medicine 32(2), 609 to 615. Graded entry.
Related SuperSkills research#
On the metric, what is the METR time horizon. On error rates, how often is AI wrong. On the mismatch between benchmarks and work, what is the jagged frontier. On when a model improves a decision, when does AI improve a decision.
About this research#
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. The two METR posts were read at source on 3 October 2026, and the other three studies were read at source on 30 September and 1 October 2026; all are graded in the evidence base. Capability slope and capability level are descriptive terms and not SuperSkills coinages.
Evidence review · SS-2026-388 · Graded against the published rubric · 2 working papers
Hirji, R. (2026). Will AI get better at this soon?. The SuperSkills evidence base, SS-2026-388. https://thesuperskills.com/research/will-ai-get-better-at-this-soon. Last reviewed 3 October 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work