A pattern is forming that nobody has named properly. A person three years into a career is asked to review AI-generated work of a kind they have never produced themselves. They sign it off, because that is the process, and because there is nothing in the output that tells them not to. The organisation records a control as satisfied. Nothing has been checked.
This is the ordinary consequence of two decisions most organisations have already made independently, rather than a hypothetical risk arriving later: automate the junior work, and keep a human in the loop.
The test that defines the problem
Could the person supervising this have produced it themselves, well enough to notice if it were wrong? Where the answer is no, supervision has become approval. The distinction is invisible in every management system currently in use, because both produce the same artefact: a name against a piece of work.
Why it is happening now rather than before
Supervision has always involved reviewing work you did not personally do. What has changed is the route by which supervisors became competent.
Traditionally you supervised work you had done badly for several years and been corrected on. The competence to review was a by-product of having produced. Remove the producing, and the by-product does not appear. That is the missing rungs reaching the point where it stops being a pipeline problem and becomes a control problem.
Two other changes make it sharper. Machine output arrives with the register and structure of competence, so there is no surface signal of error, which is the subject of why AI sounds so confident. And volume rises, so review time per item falls, and that is the condition under which automation complacency worsens.
What the evidence supports, and what it does not
The mechanism is well established even though this specific configuration has not been studied.
Bainbridge described the shape in 1983: automating the routine leaves the human with the hardest residue, monitoring, while removing the practice that built the competence for it. Vaccaro and colleagues found human-AI combinations underperforming the better party alone, with losses concentrated in decision tasks where a person judges whether a system is right. Yu and colleagues found the effect of AI assistance on radiologists running from strongly positive to strongly negative between individuals, unpredicted by experience.
What does not exist is a study measuring supervision quality as a function of the supervisor's ability to perform the underlying task. Nobody has run it. This page is therefore an argument from established mechanisms rather than a finding, and it should be read as such.
Why organisations cannot see it
Three reasons, and they compound.
The output is fine. Most AI-generated work is correct, so an unqualified reviewer approving it produces good outcomes almost all the time. The failure surfaces only on the rare wrong item, which is exactly when the review mattered.
Nobody is asked the question. No process anywhere asks a reviewer whether they could have done the work. Asking it privately, once, is the cheapest diagnostic available and it is almost never run.
Admitting it is career-damaging. A person whose role is to review is not incentivised to report that they cannot. Silence here is rational, which means the absence of complaints evidences nothing about competence.
It is now a compliance exposure, not only a management one
Article 14 of the EU AI Act, in force since 2 August 2026, requires that people assigned to oversee high-risk systems are enabled to understand the system's capacities and limitations well enough to detect anomalies, to remain aware of automation bias, to interpret output correctly, and to disregard or override it.
Every one of those is a capability claim about a specific person. An organisation whose supervisors could not produce the work they review cannot substantiate any of them, and its documentation will say the control is in place. See meaningful human oversight. This is not legal advice.
Why training will not fix this
The instinct is to fix this with training, and training will not fix it. The competence in question is judgement built from repetitions, which a course cannot deliver and which the organisation has stopped producing.
There are only three honest responses, and the first is the one nobody chooses.
Restore the reps. Deliberately keep people doing enough of the underlying work to stay able to judge it. This costs real money and looks like paying people to do something a machine does faster. It is actually paying for the ability to notice when the machine is wrong. It is the only option that preserves the control.
Move the supervision. Assign review to someone who genuinely retains the competence, and accept that this is a smaller and more expensive group than the current process assumes.
Declare the gap. Record that this work is not meaningfully reviewed, set the consequence tier accordingly, and stop claiming a control you do not have. An honest gap can be managed. An assumed control cannot, and it fails at the worst possible moment.
What to do this quarter
- Ask every named reviewer, privately, whether they could produce the work. The answers will be uncomfortable and they are the finding.
- Count overrides by reviewer. A reviewer who has never disagreed with the system is either supervising nothing or supervising something they cannot assess.
- Fill in the Delegation Boundary Map for one real process. The verification column and the capability test make this visible in about ninety minutes.
- Stop describing sign-off as verification. Sign-off accepts accountability for an outcome. Verification establishes whether the content is correct. Conflating them is how the gap stays hidden.
Related SuperSkills research
On the cause, the missing rungs and synthetic seniority. On ownership, who owns verification. On why the loop fails, human in the loop is not a safeguard. On the legal duty, meaningful human oversight. On the underlying erosion, capability debt.
Key sources
- Bainbridge, L. (1983). Ironies of Automation. Automatica, 19(6).
- Vaccaro, M., Almaatouq, A. and Malone, T. (2024). When combinations of humans and AI are useful.
- Yu, F. et al. (2024). Heterogeneity and predictors of the effects of AI assistance on radiologists.
- Parasuraman, R. and Manzey, D. H. (2010). Complacency and Bias in Human Use of Automation.
- Article 14, Human Oversight, Regulation (EU) 2024/1689.
About this research
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026) and founder of The SuperSkills Intelligence Company. This is an argument from established mechanisms rather than a finding: the specific configuration has not been studied, and the page says so above. Not legal advice. Reviewed quarterly.
Cite this
Hirji, R. (2026). Who supervises work they cannot do themselves? The SuperSkills Intelligence Company. Last reviewed 26 August 2026. thesuperskills.com/research/who-supervises-work-they-cannot-do