Production capacity inside organisations has moved a long way in three years. Checking capacity has not moved at all, because checking capacity is a quantity of senior judgement and no tool manufactures that.
A firm can responsibly release only as much work as it can check. Treat that as an arithmetic constraint rather than a cultural complaint and it changes what a leadership team is deciding. The question stops being how much more the team can produce and becomes how much more it can stand behind in a week.
The two curves are not parallel#
METR's benchmark work puts the 50 per cent task-completion horizon, the length of expert task a model finishes about half the time, as doubling every 207 days from 2019 to 2025, with a confidence interval of 166 to 240 days. The 80 per cent horizon doubles on a similar clock and sits well below it. The authors state the tasks are systematically different from real work: automatically scored, no other agents involved, few real consequences. A January 2026 re-estimate on an expanded suite of 228 tasks put the full-period doubling at 196 days, with faster recent trends of 131 days since 2023 and 89 days since 2024, and METR notes that the trend is somewhat sensitive to task composition and that the intervals remain very wide.
Against that, the quantity of experienced judgement inside a firm grows at the rate people grow, which is years. Nothing in the adoption of these tools adds to it, and the pages on the missing rungs and capability debt describe the mechanisms by which it quietly contracts.
Two curves on different clocks produce a backlog. The backlog is invisible in the only metrics most organisations have, because it shows up as work that is finished and waiting rather than as work that was not done.
What happens when checking is the residue#
Bainbridge's analysis of automated process control is from 1983 and remains the most useful statement of the problem. Automating the routine parts of a task leaves the human with the hardest residue, monitoring and exception handling, while removing the routine practice that built the competence to do it. The job gets harder while looking easier. Her argument is about process control and transfers to generative systems by analogy rather than by measurement.
Parasuraman and Manzey, reviewing automation bias and complacency across aviation, medicine and military settings, found under-questioning of automated advice in novices and experts alike, resistant to training, and worse under workload. The effect size for generative systems is less predictable than for the automation they studied. The workload finding is the one that bears on capacity: pushing more output through the same checkers does not produce proportionally more checking. It produces faster approval.
Vaccaro, Almaatouq and Malone found human and AI combinations performing worse than the better of the two alone across 106 experiments, with losses concentrated in decision tasks. Adding a reviewer is not automatically a control, and an under-resourced reviewer may be worse than none.
The field has built watching and not stopping#
Shi and DiFranzo audited 63 public artefacts on agent systems and counted which accountability controls were described. Monitoring and tool mediation appeared in about 40. Checkpoint placement appeared in 6, validator independence in 4, recovery in 2, contestability in 1. Their conclusion is that observability can become a substitute for accountability when it shifts verification onto users after meaningful intervention is no longer possible.
It is a preprint counting public descriptions rather than deployed systems. Read as a map of attention it says that the industry has concentrated on seeing what happened and has barely addressed independent checking, undoing or appeal, which are the three things a checking function needs to do its job.
Lee and colleagues, surveying 319 knowledge workers about 936 real uses of AI at work, found the character of professional thinking shifting from producing to verifying and from solving to integrating, with higher confidence in the tool associated with less critical thinking. It is self-report and cannot separate cause from selection. It describes a workforce whose job has already moved towards the constraint.
Running a firm to its checking capacity#
- Set a release rate and own it. Decide how much goes out per week on the basis of what can be checked, and say so, rather than discovering the limit through an incident.
- Count the catches. Organisations count output and not errors caught. A log of what the check found is the only local estimate of how much checking the work needs.
- Apply the capability test before the rate. Could this reviewer have produced this work well enough to notice if it were wrong? Where the answer is no, the control is absent whatever the process diagram says, which is the subject of who supervises work they cannot do themselves.
- Protect the people who are the constraint. If four people can check, the firm's capacity is those four, and filling their week with production is spending the constraint on the cheap thing.
- Build checkers deliberately. Checking capability comes from having done the work. A firm that removes the junior production and expects senior checkers to appear later has decided to run out.
- Separate sign-off from verification. Accepting accountability for an outcome and establishing that the content is correct are different acts, and a senior person can do the first without being able to do the second.
What would settle it#
Checking time per unit of output, measured in firms, before and after delegation, with the rework counted. That number would convert this argument into a capacity model any leadership team could run. Nobody publishes it. The evidence above establishes that monitoring is demanding, that automation bias worsens under workload, and that production horizons are growing; the capacity constraint drawn from them is an argument on this site.
Where this sits in my own argument#
The verifier's discount is about how checking is priced. This is about how much of it exists. A firm that has priced verification as administration has also, without deciding to, set a ceiling on what it can safely ship, and that ceiling falls as the people who could check retire or leave. Choosing the release rate deliberately is the form drift versus design takes inside an operating plan.
Related SuperSkills research#
On pricing the work of checking, the verifier's discount. On review positions that cannot function, human in the loop is not a safeguard. On setting the requirement stage by stage, the delegation boundary map. On who answers for it, who owns verification when AI does the work?
Key sources
- Kwa, T. et al., METR (2025). Measuring AI Ability to Complete Long Software Tasks.
- METR (2026). Time Horizon 1.1.
- Bainbridge, L. (1983). Ironies of Automation.
- Parasuraman, R. and Manzey, D. H. (2010). Complacency and Bias in Human Use of Automation.
- Vaccaro, M., Almaatouq, A. and Malone, T. (2024). When combinations of humans and AI are useful.
- Shi, H. and DiFranzo, D. (2026). When Agents Act Unwatched.
- Lee, H.-P. et al. (2025). The Impact of Generative AI on Critical Thinking.
About this research#
Written by Rahim Hirji, author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company.
How this research works · Reviewed quarterly · Found an error? Tell me and it is corrected on the page.
Evidence review · SS-2026-390 · Graded against the published rubric · 5 peer-reviewed studies and 2 working papers
Hirji, R. (2026). How Much Can One Firm Check?. The SuperSkills evidence base, SS-2026-390. https://thesuperskills.com/research/how-much-can-one-firm-check. Last reviewed 3 October 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work