← Research
Research

Shallow jobs

Asking humans to review every output does not scale, and it produces the appearance of oversight rather than oversight.

Last reviewed: 5 September 2026

Shallow jobs is Bain's term, from a brief of 4 June 2026. This page defines it, explains the mechanism that produces it, and separates it from the human-in-the-loop prescription it is usually mistaken for.

Question this page answersAll 811 questions this research covers

Shallow jobs are roles in which people rubber-stamp mostly correct AI output without engaging their judgement. Bain named them in a brief published on 4 June 2026, reporting what financial services executives concluded at a summit on AI. Bain's emphasis falls on mostly correct: the output is usually right, and being usually right is what stops the reviewing.

The answer, in one line

Shallow jobs are roles in which people rubber-stamp mostly correct AI output without engaging their judgement. Bain named them in June 2026, reporting on a summit of financial services executives, and the emphasis falls on mostly correct: the output is usually right, which is what makes the reviewing stop.

Share as a card

The mechanism#

A shallow job is produced by a decision that sounds responsible: require a human to review every output. Bain's example is payments, a process running at volumes where universal review cannot be thoughtful. The reviewer faces a queue too large to consider properly, containing items that are almost always fine. Approving becomes the rational response to the workload, and the role settles into a rhythm of confirmation.

The failure belongs to the design rather than the person. A reviewer given ten items a day can think about each one. A reviewer given four hundred cannot, and no amount of diligence changes the arithmetic. What the organisation has built is a checkpoint that reports supervision while performing throughput.

Human-factors research explains why this happens so reliably. Mackworth showed in 1948 that detection accuracy for rare signals had already fallen measurably by the end of the first half hour of a two-hour watch, and kept falling. He analysed in half-hour blocks, so the decline cannot be placed more precisely than that. The effect has survived seventy years of replication, though the argument about its mechanism has not settled. Automation complacency describes the same drift towards trust after a history of reliable performance. A process that is right ninety-nine times in a hundred trains the reviewer to expect the ninety-nine.

Not the same as human in the loop, and often its result#

Shallow jobs are what the human-in-the-loop prescription becomes when it is applied without limit. Narayanan and Kapoor reach the same conclusion from a different direction, arguing that a system requiring review and approval of every AI decision either devolves into the human acting as a rubber stamp or is outcompeted by a less safe solution that does not.

That is a sharper claim than it first appears. Blanket review carries the cost of oversight without the benefit, so it competes badly against systems that skip it altogether. The organisation pays for supervision it is not receiving, and the arrangement is unstable for that reason.

The distinction the European vocabulary draws is useful here. Human in the loop describes where somebody sits. Human in command, the term the European Economic and Social Committee uses, describes what they are entitled and able to do. A shallow job satisfies the first and fails the second.

The compliance exposure#

Rubber-stamping is the specific behaviour regulators test for, which makes this more than a design preference.

Under UK GDPR Article 22A, a decision is solely automated where there is no meaningful human involvement. The Information Commissioner's Office sets out the test in its recruitment work: whether a human can exercise real influence over a decision before it is applied, and holds the authority, discretion and competence to alter it. Competence is written into the test. A reviewer approving a queue holds none of the three, so a process staffed by shallow jobs may be legally automated while every organisation chart shows a person in place.

The same logic sits inside the EU AI Act, whose high-risk tier requires oversight that is meaningful rather than nominal. An organisation that has allowed the underlying competence to decay cannot restore it with a sign-off box.

What to design instead#

Bain's alternative is to build roles around the things people do that machines do not: judgement on the cases the system cannot resolve, training the system, accountability for outcomes, and empathy and trust where those matter. Their framing is that this is a workforce-design choice made deliberately rather than a backstop bolted onto an automated process.

Read against the rest of this research, that translates into a sequence. Decide which cases require a person before the process is built, rather than reviewing everything afterwards. Keep the reviewer's own practice alive, because verification depends on competence the reviewer must still possess. And accept that targeted judgement on a small number of hard cases delivers more oversight than nominal review of everything.

Reducing the amount reviewed can therefore increase the amount of supervision. Volume of checking and quality of checking pull against each other once the queue exceeds what attention can carry.

Where Bain disagrees with itself#

Anyone citing the firm on this should know it argues both sides. Six weeks before the shallow jobs brief, Bain published an argument that juniors learn by reviewing, stress-testing and catching errors in AI-generated output, with the repetitions per hour going up rather than down. Learning by verifying and shallow jobs describe the same activity, reaching opposite conclusions about what it does to a person.

Both can hold. Reviewing thirty drafted models a week builds instinct when somebody senior is examining the review, which is the medical residency structure Bain invokes. The same thirty models produce a shallow job when nobody is. The variable is whether anyone is watching the reviewer.

Key sources

Explainer · SS-2026-180 · Graded against the published rubric

Cite this page

Hirji, R. (2026). Shallow jobs. The SuperSkills evidence base, SS-2026-180. https://thesuperskills.com/research/what-are-shallow-jobs. Last reviewed 5 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

What are shallow jobs?

Shallow jobs are roles in which people rubber-stamp mostly correct AI output without engaging their judgement. Bain named them in June 2026, reporting on a summit of financial services executives, and the emphasis falls on mostly correct: the output is usually right, which is what makes the reviewing stop.

What causes a shallow job?

Requiring human review of every output in a high-volume process. Bain's example is payments. Review designed as universal coverage rather than targeted judgement produces a queue too large to think about, and the reviewer adapts by approving. The failure is in the workflow design rather than in the person.

Is a shallow job the same as human in the loop?

A shallow job is what human in the loop becomes when it is applied without limit. Narayanan and Kapoor make the same argument, describing a system requiring review of every decision as one that either devolves into the human acting as a rubber stamp or is outcompeted by a less safe solution. Being positioned in a process is not the same as exercising authority over it.

What should replace blanket review?

Bain argues for designing roles around what humans are uniquely good at: exercising judgement on the cases AI cannot resolve, training the system, being accountable for outcomes, and bringing empathy and trust where they matter. Their framing is that this is a workforce-design choice made deliberately, rather than a backstop bolted onto an automated process.

Why do shallow jobs create a compliance risk?

Because rubber-stamping is the specific thing regulators test for. Under UK GDPR Article 22A a decision counts as solely automated unless a human can exercise real influence before it is applied and has the authority, discretion and competence to alter it. A reviewer approving a queue meets none of those conditions, so a process staffed by shallow jobs may be legally automated while appearing supervised.

In this hub

Definitions

The terms this field uses, defined against their primary sources.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Oversight is the topic most often agreed with in principle and least often implemented. There is the human oversight version, and the full range of topics and audiences.

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.