← Research
Research

Who owns verification when AI does the work?

In most organisations, nobody. It is not in a job description, a budget line or an org chart.

Last reviewed: 26 August 2026

Why verification disappears when generation is separated from the person, the one test that settles whether it is real, four ownership models with their actual costs, and why verification capacity depreciates exactly when you automate.

Questions this page answersAll 780 questions this research covers

In most organisations, nobody. Verification is the one stage of AI-assisted work that everybody assumes is happening and almost nobody has assigned. It is not in a job description, not in a budget line, not on an org chart, and not usually in the process document. It is assumed to be a property of the workflow rather than the responsibility of a person, and assumed responsibilities are the ones that fail .

This is now a live legal question as well as a management one. Since 2 August 2026, Article 14 of the EU AI Act requires that people assigned to oversee a high-risk system are enabled to detect anomalies, interpret output correctly, and disregard or override it. That presupposes there is somebody assigned. Many organisations will discover, under audit, that there is not.

Why it disappears#

It was never a job, it was a by-product. When a person wrote the report, checking was part of writing it. Separate the generation from the person, and the checking becomes a distinct task nobody was hired to do and nobody has time allocated for. It falls into the gap between the person who prompted and the person who signed.

It is priced as administration. Producing is visible, senior and rewarded. Checking is invisible until it fails, and the person who catches an error gets less credit than the person who produced the output that contained it. That mispricing is what I call the verifier's discount, and it reliably produces less verification than an organisation thinks it has bought.

Sign-off gets mistaken for verification. They are different acts. Sign-off is accepting accountability for an outcome. Verification is establishing whether the content is correct. A senior person can perform the first without being able to perform the second, and in AI-assisted work that combination is now common.

The people best placed to do it are the ones being removed. The junior work that produced the pattern recognition needed to spot a wrong answer is the work most easily automated. See the missing rungs.

The test that settles it#

The capability test#

Could the person verifying this have produced it themselves, well enough to notice if it were wrong?

If no, the verification is decorative. It should be recorded as absent rather than as satisfied, because recording it as satisfied is what turns a capability gap into a governance failure. This is the only test on this page that cannot be gamed. It is the one almost nobody applies.

The reason it is decisive is that fluency carries no signal. A language model's confidence is a property of its writing style rather than of its knowledge, so a plausible wrong answer and a correct one look identical to anyone who cannot independently evaluate the content. Reviewing without competence is a different activity that produces a similar-looking record.

Four ownership models, and what each actually costs#

The producer verifies. Whoever prompted checks their own output. Cheapest, and it fails on exactly the errors that matter, because someone who accepted a framing is poorly placed to notice the framing was wrong. Workable for low-consequence work, not for anything else.

A named peer verifies. Someone with equivalent domain competence, not in the production chain. This is the model that passes the capability test most often. It costs real time and it is the first thing cut when throughput matters.

A specialist function verifies. A dedicated group, as model risk management works in banking, which is the most mature precedent available and worth studying rather than reinventing. Strong for high-consequence, repeated decisions. Expensive, and it can turn into a rubber stamp if the function lacks authority to stop work.

Nobody verifies, declared. For low-stakes output, this is a legitimate and honest choice, and far better than pretending. The failure mode is not choosing this; it is choosing it by default and describing it as one of the other three.

The four models have not been compared#

No study has compared these four models for accuracy or cost. The recommendation to prefer a named peer rests on the capability test and on the Vaccaro meta-analysis finding that human-AI combinations underperform the better party when the human is judging rather than producing, not on a trial of ownership structures. It is a reasoned position rather than a demonstrated one, and should be read that way.

There is also a real objection. Where a system genuinely outperforms every available human verifier, insisting on human verification buys accountability at the price of accuracy. That may be the right trade for legitimacy and for the ability to explain a decision, but it is a trade, and organisations should make it deliberately rather than assume they are getting both.

Verification is expertise, applied#

Verification is not a separate activity from expertise. It is expertise, applied. That single reframing changes what an organisation does about it.

If verification is administration, you buy more of it cheaply and treat it as overhead. If verification is expertise applied, then verification capacity is a stock that has to be maintained, and it depreciates precisely when you automate the work that built it. An organisation that automates production while assuming verification will look after itself is spending down an asset it never put on the balance sheet. That is capability debt in its most operational form.

Which produces an uncomfortable rule: you cannot automate a task and retain the ability to verify it unless you deliberately fund the practice that keeps someone competent at it. That looks like paying people to do work a machine does faster, so almost nobody does it. It is actually paying for the ability to notice when the machine is wrong, and after 2 August 2026 in the EU it is also paying for the ability to pass an audit.

What to do this quarter#

On the mispricing, the verifier's discount. On the legal duty, meaningful human oversight. On the practical tool, the Delegation Boundary Map. On detecting error at all, how do I know when AI is wrong and automation bias. On the underlying erosion, capability debt. The position that follows from this, put simply, is that human in the loop is not a safeguard. See who supervises work they cannot do.

Key sources

About this research#

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. The regulatory position is quoted from the primary text and dated; the interpretation is the author's and kept separate. Not legal advice.

How this research works  ·  Reviewed quarterly  ·  Found an error? Tell me and it is corrected on the page.

Cite this

Hirji, R. (2026). Who owns verification when AI does the work? The SuperSkills Intelligence Company. Last reviewed 26 August 2026. thesuperskills.com/research/who-owns-verification-when-ai-does-the-work

Questions answered on this page

Who owns verification when AI does the work?

In most organisations, nobody. Verification is not in a job description, a budget line, an org chart or usually the process document. It is assumed to be a property of the workflow rather than the responsibility of a named person. Since 2 August 2026, Article 14 of the EU AI Act requires that people assigned to oversee a high-risk system can detect anomalies and override output, which presupposes somebody is assigned. Many organisations will find under audit that nobody is.

How do you tell whether verification is real or decorative?

Apply the capability test: could the person verifying this have produced it themselves, well enough to notice if it were wrong? If no, the verification is decorative and should be recorded as absent rather than as satisfied. The test is decisive because fluency carries no signal: a model's confidence is a property of its writing style rather than its knowledge, so a plausible wrong answer and a correct one look identical to anyone who cannot independently evaluate the content.

Is sign-off the same as verification?

No. Sign-off is accepting accountability for an outcome. Verification is establishing whether the content is correct. A senior person can perform the first without being able to perform the second, and in AI-assisted work that combination is now common. Treating a signature as evidence of checking is one of the main ways verification disappears without anyone noticing.

What are the options for assigning verification?

Four. The producer verifies, which is cheapest and fails on framing errors. A named peer with equivalent domain competence verifies, which passes the capability test most often and is the first thing cut under throughput pressure. A specialist function verifies, as model risk management works in banking, which suits high-consequence repeated decisions but can become a rubber stamp without authority to stop work. Or nobody verifies and this is declared, which is legitimate for low-stakes work and far better than pretending. The failure mode is choosing the fourth by default and describing it as one of the others.

In this hub

Judgement, oversight and accountability

Who decides, who checks, and who is answerable when the machine was involved.

Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Verification with a named owner who survives that person leaving is rare. Assigning it properly, so it holds when somebody resigns, is what I come in to do. AI advisory for CEOs and boards.

Oversight is the topic most often agreed with in principle and least often implemented. There is the human oversight version, and the full range of topics and audiences.

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.