Every organisation now knows what its AI tools cost per seat. Very few know what delegation costs, which is a different number and a larger one. It includes the time spent briefing, the time spent checking, the rework when the check finds something, the occasions when the work was handed over and should not have been, and the difference between a task feeling quicker and being quicker.
None of those sit in a procurement line. All of them land on named people, usually the most experienced ones, and the measurement problem is the reason most organisations cannot say whether the delegation is paying.
Speed and value are different sizes, measured on the same people#
METR surveyed 349 technical workers between February and April 2026: software engineers, researchers, academics, founders and managers. The median self-reported change in the value of their work was between 1.4 and 2 times. The median self-reported change in speed was 3 times. Asked about different years, the same respondents placed themselves at 1.3 times in March 2025, 2 times in March 2026, and forecast 2.5 times for 2027.
Every number there is self-reported, from a convenience sample with a response rate of about 2 per cent, drawn from people inclined to volunteer for a survey about AI. It cannot establish a productivity effect and the authors do not claim one. What it does carry is the gap: on the same respondents, in the same instrument, speed runs at roughly double value. An organisation measuring only the first will record a gain twice the size of the one it is getting.
People are specifically miscalibrated about this#
Yu and colleagues ran a preregistered study with 1,237 participants on short cognitive tasks, comparing forecast completion times against actual ones. Actual completion times did not differ between independent and AI-assisted completion, while participants predicted that AI would be significantly faster. The same bias did not appear when participants imagined help from another person, and they reported lower subjective effort with AI at equivalent completion times.
The tasks were short and simple by design, so this says nothing directly about professional work or output quality. The control condition is the interesting part. The miscalibration is about machine help in particular rather than about assistance in general, and lower effort was being experienced as less elapsed time.
What can be handed over is moving, and not evenly#
METR's benchmark work on task length puts the 50 per cent completion horizon, meaning the length of expert task a model finishes about half the time, as doubling every 207 days from 2019 to 2025, with a bootstrapped confidence interval of 166 to 240 days. The 80 per cent horizon doubles on a similar clock but sits well below the 50 per cent figure.
That pair of numbers is the most useful thing in this base for anyone planning delegation. The headline horizon describes what a model will attempt at a coin-flip success rate. The reliability most professional work actually needs sits on the lower line. The authors also state that the tasks are systematically different from real work: all are automatically scored, none involves other agents, and few carry real consequences.
METR re-estimated the measure in January 2026 on an expanded suite of 228 tasks, up from 170, and put the full-period doubling at 196 days, which it describes as unchanged from its own re-derived baseline rather than from the 207 published earlier. Recent trends came out faster: 131 days since 2023 and 89 days since 2024. Its own cautions belong with the figures: the trend is somewhat sensitive to how the task suite is composed, the confidence intervals remain very wide, and human baseline times exist for only 5 of the 31 longest tasks.
Undesigned pairing can subtract#
Vaccaro, Almaatouq and Malone ran a preregistered meta-analysis of 106 experimental studies and 370 effect sizes and found human and AI combinations performing significantly worse on average than the better of human or AI alone, at a Hedges' g of minus 0.23. Losses concentrated in decision-making; gains appeared in content creation. Pairing gained where the human was stronger and lost where the model was.
The benchmark is an oracle-selected best performer, which an organisation rarely knows in advance, and the studies predate current frontier models. The direction still matters for costing: adding a person to a machine is not free and is not automatically a control.
And the macro number is small#
Acemoglu's task-based model estimates total factor productivity gains of no more than 0.66 per cent over ten years, revised to under 0.53 per cent once hard-to-learn tasks are accounted for. The model works through task-level cost savings and would not capture effects running through new products or new tasks, so it is a floor on one mechanism rather than a ceiling on all of them. Anyone building a business case on an order-of-magnitude larger figure should be able to say which mechanism they are claiming.
How to cost it#
- Put the checking in the plan with a name against it. Verification with no allocated time does not happen under pressure, and pressure is when it matters. The pricing of that work is the verifier's discount.
- Measure value, not speed. Ask what changed about the decision or the deliverable, not how long it took. The two diverge by a factor of two in the best available self-report.
- Record what came back wrong. Organisations count output and not catches, so one is visible and the other is not. A log of what the check found is the only source of a local reliability estimate.
- Match the handover to the stakes, not to the capability. Keep, share, hand over sets the default conservatively, and the delegation boundary map sets the verification requirement stage by stage.
- Decide where the freed time goes before it arrives. Otherwise it goes nowhere anybody chose, which is the unclaimed hour.
- Keep the receipts. Hand, Head, Hours, Heart is the short form of what a human contribution looked like, kept where somebody can find it later.
What would settle it#
Elapsed time and output quality on real professional tasks, measured rather than reported, in firms, before and after delegation, with the checking time counted. METR's own authors are explicit that their survey is not that study. Nobody has published it, so every cost estimate in circulation, including the ones on this page, rests on self-report or on benchmarks the authors say differ from real work.
Where this sits in my own argument#
Measuring use rather than value is usage theatre, and it survives because the honest measurement is harder. The cost of delegation is where the argument on this site meets a finance function: not whether the capability is real, but what a firm spent to get it and who paid in attention.
Related SuperSkills research#
On the unmeasured capacity, the unclaimed hour. On measuring the wrong thing, usage theatre. On the price of checking, the verifier's discount. On what to hand over at all, keep, share, hand over.
Key sources
- Becker, J., METR (2026). Measuring the Self-Reported Impact of Early-2026 AI on Technical Worker Productivity.
- Yu, S. et al. (2026). Cognitive offloading and the speedup illusion in human-AI interaction.
- Kwa, T. et al., METR (2025). Measuring AI Ability to Complete Long Software Tasks.
- METR (2026). Time Horizon 1.1.
- Vaccaro, M., Almaatouq, A. and Malone, T. (2024). When combinations of humans and AI are useful.
- Acemoglu, D. (2024). The Simple Macroeconomics of AI.
About this research#
Written by Rahim Hirji, author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company.
How this research works · Reviewed quarterly · Found an error? Tell me and it is corrected on the page.
Evidence review · SS-2026-392 · Graded against the published rubric · 4 peer-reviewed studies, 1 working paper and 1 institutional survey
Hirji, R. (2026). The Cost of Delegation. The SuperSkills evidence base, SS-2026-392. https://thesuperskills.com/research/the-cost-of-delegation. Last reviewed 3 October 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work