"Keep a human in the loop" is the most repeated instruction in organisational AI, and it is close to useless as guidance. It does not say which human, with which capability, at which point in the work, or on what grounds they may say no. Left unspecified, it collapses into a person approving output they did not produce, under time pressure, in a domain they may not be able to check. That is not oversight. It is a signature.
The Delegation Boundary Map replaces the instruction with a decision. It breaks work into nine stages, and for each one asks four questions: what does the human retain, what may AI assist with, what may AI execute alone, and who owns the outcome if it goes wrong. Filling it in takes a team about ninety minutes. The argument it usually produces is the point.
Why nine stages rather than a loop
A loop implies the human is positioned somewhere in a circle. Work is not a circle. It is a sequence in which authority can be handed over at any point, and the consequences of handing it over are completely different at the start than at the end. Delegating problem definition and delegating execution are not the same act and should never be governed by the same rule.
The evidence this rests on
Three findings make an explicit boundary necessary rather than merely tidy.
Combination does not reliably beat the better half. A meta-analysis of 370 effect sizes across 106 experiments found human-AI combinations performed worse on average than the stronger of human alone or AI alone. Losses concentrated in decision tasks, where the human judges whether the system is right. Gains appeared in creation tasks, where the pair produces something together. A map that moves human involvement earlier converts the first into the second.
Averages conceal opposite effects. The effect of AI assistance on radiologists ran from strongly positive to strongly negative between individuals, and was not predicted by experience or prior familiarity with AI. A blanket rule of the form "clinicians will use the tool" is a rule with unknown sign. Boundaries have to be set per stage and checked per person.
Monitoring is the hardest residue, not the easiest. Bainbridge set this out in 1983: automate the routine and you leave the human with the task that requires the most skill, while removing the practice that built it. Forty years and one technology later, nothing about that has been repealed.
The nine stages
Work the sequence in order. The stages before delegation are where most of the value sits and where almost nobody looks.
1 · Problem definition
What is actually being solved, and is it the right problem? Human retains this entirely. A model asked to solve the wrong problem will solve it fluently, and fluency is exactly what makes a misframed problem hard to catch downstream. Accountability: the person who owns the outcome.
2 · Intent
What is this work for, and what would count as success? Human retains. AI may assist by surfacing options you had not considered, but only after your own intent is written down. Reverse that order and you have anchored yourself to the machine's framing before forming one of your own.
3 · Context
What does the system not know? History, politics, relationships, the last three times this was tried. Human supplies; AI may organise. This stage is where most AI-assisted work quietly fails, because context is invisible in the output and its absence looks like confidence.
4 · Constraints
Budget, regulation, risk appetite, what is off the table, and the rejection criteria. Human sets, and sets them before generation. Deciding what would make you say no is far easier before an answer is in front of you looking finished.
5 · Delegation
The decision itself: what is being handed over, and on what terms. Human decides, explicitly and in writing. This is the stage almost every organisation skips, which is why delegation happens by accumulation rather than by choice. Accountability for the delegation decision sits with the person making it, not with whoever later uses the output.
6 · Generation and execution
AI may execute, within the constraints set at stage four. This is the stage everyone argues about and it is the least interesting one, because if stages one to five were done properly the risk here is bounded, and if they were not, no amount of review at stage seven will recover it.
7 · Verification
Human retains, and this is the stage with the hard requirement. The verifier must be capable of detecting the error. That single test governs everything: if nobody in the chain could produce the work themselves, the verification is decorative and should be recorded as absent rather than as satisfied.
Never verify against another model. Verify against a different kind of source: a primary document, a person who knows, the original study. The South African High Court judgment in Mavundla records a judge testing a suspect citation in ChatGPT, which confirmed the fabricated case was real, and then confirmed it addressed a point it could not have addressed.
8 · Decision
Human, with the override rule stated in advance. On what grounds may the human overrule the system, and on what grounds must they defer? Unspecified, this collapses into whoever is more confident in the room, which is the worst available decision procedure. Where the system genuinely outperforms the human, saying so and deferring deliberately is more honest than a veto nobody exercises.
9 · Outcome ownership
Human, named, before the work starts. Not the team, not the function, a person. Accountability assigned after an outcome is not accountability, it is attribution, and the two behave completely differently under pressure.
The four tests that make it real
- The capability test. Could the person verifying this detect the error? If not, the boundary is in the wrong place. This is the test that connects the map to capability debt: every boundary you set has a maintenance cost, because the human half has to stay good enough to disagree.
- The baseline test. Does this arrangement beat the better of human alone and system alone? Most deployments have never measured it, which means they do not know whether the pairing is adding or subtracting.
- The naming test. Can you name the accountable person for every stage, today, without a meeting? Gaps in that list are where accountability will fail, and they are visible in advance.
- The reversal test. If this goes wrong, can you reconstruct why the decision was made? If the reasoning chain lives only inside a model's output, it has already vanished.
Where this fits
The map is the operational form of Human at the Start. That page argues the principle; this one is the artefact you fill in. It is also the answer to the objection raised in what is human-AI collaboration, that the common configuration underperforms: moving the human to stages one through five is precisely how you convert a losing decision task into a winning creation task.
A finished map is not a policy document. It is a record of decisions that were previously being made by accumulation, which is the distinction at the centre of drift versus design.
Where this is uncertain
Two honest limits. First, the nine stages are a working framework drawn from practitioner research across more than two hundred organisations, not an experimentally validated model. The evidence supports the individual claims about automation bias, combination effects and monitoring; it does not yet validate this particular sequence against alternatives. Second, no study has tested whether organisations that set explicit boundaries outperform those that do not. That would be a genuinely useful experiment and nobody has run it.
The map is therefore offered as a decision-forcing device rather than as a proven intervention, and it is described that way deliberately.
Use it
The working version is a nine-row grid: stage, human retains, AI may assist, AI may execute, verification requirement, accountability owner. Take it into a room with the people who actually do the work, not only the people who own the budget, and fill it in for one real process rather than in the abstract.
Expect disagreement at stages five and seven. That disagreement is the output. A map everybody agreed with immediately was filled in by one person describing what already happens.
Download the working grid (CSV) · free to use and adapt with attribution.
Related SuperSkills research
On the principle, Human at the Start. On decision architecture, human and AI decision making and decision quality. On agents, AI agents and human judgement. On the tendency it guards against, automation bias and how do I know when AI is wrong. On why verification is under-resourced, the verifier's discount. On the organisational stakes, AI workforce strategy. On the legal requirement this now sits alongside, meaningful human oversight. On assigning the verification stage specifically, who owns verification.
Key research and primary sources
- Vaccaro, M., Almaatouq, A. and Malone, T. (2024). When combinations of humans and AI are useful: a systematic review and meta-analysis. Nature Human Behaviour, 8.
- Bainbridge, L. (1983). Ironies of Automation. Automatica, 19(6).
- Yu, F. et al. (2024). Heterogeneity and predictors of the effects of AI assistance on radiologists. Nature Medicine, 30(3).
- Parasuraman, R. and Manzey, D. H. (2010). Complacency and Bias in Human Use of Automation. Human Factors, 52(3).
- Mavundla v MEC: COGTA KwaZulu-Natal [2025] ZAKZPHC 2. High Court of South Africa, Pietermaritzburg.
About this framework
The Delegation Boundary Map was developed by Rahim Hirji, author of SuperSkills: The Seven Human Skills for the Age of AI (Kogan Page, 2026) and founder of The SuperSkills Intelligence Company. Developed, not coined from nothing: it builds on established work in human factors and automation, and on the practitioner research base. The findings cited are attributed to the studies that produced them. Free to use and adapt with attribution. Reviewed quarterly.
Cite this
Hirji, R. (2026). The Delegation Boundary Map. The SuperSkills Intelligence Company. Last reviewed 26 August 2026. thesuperskills.com/research/delegation-boundary-map