Shallow jobs are roles in which people rubber-stamp mostly correct AI output without engaging their judgement. Bain named them in a brief published on 4 June 2026, reporting what financial services executives concluded at a summit on AI. Bain's emphasis falls on mostly correct: the output is usually right, and being usually right is what stops the reviewing.
The answer, in one line
Shallow jobs are roles in which people rubber-stamp mostly correct AI output without engaging their judgement. Bain named them in June 2026, reporting on a summit of financial services executives, and the emphasis falls on mostly correct: the output is usually right, which is what makes the reviewing stop.
The mechanism#
A shallow job is produced by a decision that sounds responsible: require a human to review every output. Bain's example is payments, a process running at volumes where universal review cannot be thoughtful. The reviewer faces a queue too large to consider properly, containing items that are almost always fine. Approving becomes the rational response to the workload, and the role settles into a rhythm of confirmation.
The failure belongs to the design rather than the person. A reviewer given ten items a day can think about each one. A reviewer given four hundred cannot, and no amount of diligence changes the arithmetic. What the organisation has built is a checkpoint that reports supervision while performing throughput.
Human-factors research explains why this happens so reliably. Mackworth showed in 1948 that detection accuracy for rare signals had already fallen measurably by the end of the first half hour of a two-hour watch, and kept falling. He analysed in half-hour blocks, so the decline cannot be placed more precisely than that. The effect has survived seventy years of replication, though the argument about its mechanism has not settled. Automation complacency describes the same drift towards trust after a history of reliable performance. A process that is right ninety-nine times in a hundred trains the reviewer to expect the ninety-nine.
Not the same as human in the loop, and often its result#
Shallow jobs are what the human-in-the-loop prescription becomes when it is applied without limit. Narayanan and Kapoor reach the same conclusion from a different direction, arguing that a system requiring review and approval of every AI decision either devolves into the human acting as a rubber stamp or is outcompeted by a less safe solution that does not.
That is a sharper claim than it first appears. Blanket review carries the cost of oversight without the benefit, so it competes badly against systems that skip it altogether. The organisation pays for supervision it is not receiving, and the arrangement is unstable for that reason.
The distinction the European vocabulary draws is useful here. Human in the loop describes where somebody sits. Human in command, the term the European Economic and Social Committee uses, describes what they are entitled and able to do. A shallow job satisfies the first and fails the second.
The compliance exposure#
Rubber-stamping is the specific behaviour regulators test for, which makes this more than a design preference.
Under UK GDPR Article 22A, a decision is solely automated where there is no meaningful human involvement. The Information Commissioner's Office sets out the test in its recruitment work: whether a human can exercise real influence over a decision before it is applied, and holds the authority, discretion and competence to alter it. Competence is written into the test. A reviewer approving a queue holds none of the three, so a process staffed by shallow jobs may be legally automated while every organisation chart shows a person in place.
The same logic sits inside the EU AI Act, whose high-risk tier requires oversight that is meaningful rather than nominal. An organisation that has allowed the underlying competence to decay cannot restore it with a sign-off box.
What to design instead#
Bain's alternative is to build roles around the things people do that machines do not: judgement on the cases the system cannot resolve, training the system, accountability for outcomes, and empathy and trust where those matter. Their framing is that this is a workforce-design choice made deliberately rather than a backstop bolted onto an automated process.
Read against the rest of this research, that translates into a sequence. Decide which cases require a person before the process is built, rather than reviewing everything afterwards. Keep the reviewer's own practice alive, because verification depends on competence the reviewer must still possess. And accept that targeted judgement on a small number of hard cases delivers more oversight than nominal review of everything.
Reducing the amount reviewed can therefore increase the amount of supervision. Volume of checking and quality of checking pull against each other once the queue exceeds what attention can carry.
Where Bain disagrees with itself#
Anyone citing the firm on this should know it argues both sides. Six weeks before the shallow jobs brief, Bain published an argument that juniors learn by reviewing, stress-testing and catching errors in AI-generated output, with the repetitions per hour going up rather than down. Learning by verifying and shallow jobs describe the same activity, reaching opposite conclusions about what it does to a person.
Both can hold. Reviewing thirty drafted models a week builds instinct when somebody senior is examining the review, which is the medical residency structure Bain invokes. The same thirty models produce a shallow job when nobody is. The variable is whether anyone is watching the reviewer.
Key sources
- Bain and Company (2026). What Financial Services Leaders Are Wrestling with on AI and Organizational Transformation. Van Dijk, L., Alves, M., Fleming, R. and Mehta, B., 4 June 2026. A brief reporting an executive summit rather than a study, and the shallow jobs passage sits under "Three open debates".
- Mackworth, N. H. (1948). The Breakdown of Vigilance during Prolonged Visual Search. Quarterly Journal of Experimental Psychology, 1(1), 6 to 21. The Clock Test. Mackworth analysed in half-hour blocks across a two-hour watch, so the decline cannot be located any more precisely than by the end of the first block.
- Klein, R. M. and Feltmate, B. B. T. (2025). The vigilance decrement: its first 75 years. Frontiers in Cognition, 4. The review establishing that the effect has held, and that its mechanism remains disputed.
Explainer · SS-2026-180 · Graded against the published rubric
Hirji, R. (2026). Shallow jobs. The SuperSkills evidence base, SS-2026-180. https://thesuperskills.com/research/what-are-shallow-jobs. Last reviewed 5 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work