← Research
Research · Essay

How Much Can One Firm Check?

An organisation can release only as much as it can check, and checking capacity is a stock rather than a budget line.

Last reviewed: 3 October 2026 · Next review due: 3 October 2027

Production capacity inside organisations has moved a long way in three years. Checking capacity has not moved at all, because checking capacity is a quantity of senior judgement and no tool manufactures that.

A firm can responsibly release only as much work as it can check. Treat that as an arithmetic constraint rather than a cultural complaint and it changes what a leadership team is deciding. The question stops being how much more the team can produce and becomes how much more it can stand behind in a week.

The two curves are not parallel#

METR's benchmark work puts the 50 per cent task-completion horizon, the length of expert task a model finishes about half the time, as doubling every 207 days from 2019 to 2025, with a confidence interval of 166 to 240 days. The 80 per cent horizon doubles on a similar clock and sits well below it. The authors state the tasks are systematically different from real work: automatically scored, no other agents involved, few real consequences. A January 2026 re-estimate on an expanded suite of 228 tasks put the full-period doubling at 196 days, with faster recent trends of 131 days since 2023 and 89 days since 2024, and METR notes that the trend is somewhat sensitive to task composition and that the intervals remain very wide.

Against that, the quantity of experienced judgement inside a firm grows at the rate people grow, which is years. Nothing in the adoption of these tools adds to it, and the pages on the missing rungs and capability debt describe the mechanisms by which it quietly contracts.

Two curves on different clocks produce a backlog. The backlog is invisible in the only metrics most organisations have, because it shows up as work that is finished and waiting rather than as work that was not done.

What happens when checking is the residue#

Bainbridge's analysis of automated process control is from 1983 and remains the most useful statement of the problem. Automating the routine parts of a task leaves the human with the hardest residue, monitoring and exception handling, while removing the routine practice that built the competence to do it. The job gets harder while looking easier. Her argument is about process control and transfers to generative systems by analogy rather than by measurement.

Parasuraman and Manzey, reviewing automation bias and complacency across aviation, medicine and military settings, found under-questioning of automated advice in novices and experts alike, resistant to training, and worse under workload. The effect size for generative systems is less predictable than for the automation they studied. The workload finding is the one that bears on capacity: pushing more output through the same checkers does not produce proportionally more checking. It produces faster approval.

Vaccaro, Almaatouq and Malone found human and AI combinations performing worse than the better of the two alone across 106 experiments, with losses concentrated in decision tasks. Adding a reviewer is not automatically a control, and an under-resourced reviewer may be worse than none.

The field has built watching and not stopping#

Shi and DiFranzo audited 63 public artefacts on agent systems and counted which accountability controls were described. Monitoring and tool mediation appeared in about 40. Checkpoint placement appeared in 6, validator independence in 4, recovery in 2, contestability in 1. Their conclusion is that observability can become a substitute for accountability when it shifts verification onto users after meaningful intervention is no longer possible.

It is a preprint counting public descriptions rather than deployed systems. Read as a map of attention it says that the industry has concentrated on seeing what happened and has barely addressed independent checking, undoing or appeal, which are the three things a checking function needs to do its job.

Lee and colleagues, surveying 319 knowledge workers about 936 real uses of AI at work, found the character of professional thinking shifting from producing to verifying and from solving to integrating, with higher confidence in the tool associated with less critical thinking. It is self-report and cannot separate cause from selection. It describes a workforce whose job has already moved towards the constraint.

Running a firm to its checking capacity#

What would settle it#

Checking time per unit of output, measured in firms, before and after delegation, with the rework counted. That number would convert this argument into a capacity model any leadership team could run. Nobody publishes it. The evidence above establishes that monitoring is demanding, that automation bias worsens under workload, and that production horizons are growing; the capacity constraint drawn from them is an argument on this site.

Where this sits in my own argument#

The verifier's discount is about how checking is priced. This is about how much of it exists. A firm that has priced verification as administration has also, without deciding to, set a ceiling on what it can safely ship, and that ceiling falls as the people who could check retire or leave. Choosing the release rate deliberately is the form drift versus design takes inside an operating plan.

On pricing the work of checking, the verifier's discount. On review positions that cannot function, human in the loop is not a safeguard. On setting the requirement stage by stage, the delegation boundary map. On who answers for it, who owns verification when AI does the work?

Key sources

About this research#

Written by Rahim Hirji, author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company.

How this research works · Reviewed quarterly · Found an error? Tell me and it is corrected on the page.

Evidence review · SS-2026-390 · Graded against the published rubric · 5 peer-reviewed studies and 2 working papers

Cite this page

Hirji, R. (2026). How Much Can One Firm Check?. The SuperSkills evidence base, SS-2026-390. https://thesuperskills.com/research/how-much-can-one-firm-check. Last reviewed 3 October 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work

In this hub

Judgement, oversight and accountability

Who decides, who checks, and who is answerable when the machine was involved.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory  ·  Boards  ·  Enquire