← Research
Research

How do humans and agents divide work across a process?

Splitting the tasks is the part that looks hard and is not. The failures are at the joins, and every discipline that has studied joins seriously is a clinical or an aeronautical one.

Last reviewed: 4 September 2026

Why function allocation lists do not survive contact with a real process, what medicine learned about discontinuities in care, what the only measured handoff intervention changed and what it did not, and where the law stops short.

Question this page answersQuestion this page partly answersAll 811 questions this research covers

You do not divide it. Seventy-five years of function allocation research says that handing each task to whichever party performs it better fails, because every assignment manufactures new work for the other party that nobody costed. The design object across a whole process is the boundary: at each point where work passes between a person and a machine, what state moves with it, whether the move is announced, whether the receiver can reconstruct the reasoning behind what they have just been handed, and who is answerable on either side. Those four questions have serious evidence behind them. None of it is about AI. It is in operating theatres, in intensive care, and in accident reports.

The answer, in one line

Not by dividing it. The evidence from fifty years of function allocation is that assigning tasks to whichever party is better at them does not produce a working system, because each assignment creates new work for the other party that did not exist before.

Share as a card

The list that has been on the wall since 1951#

The standard answer to this question is a list. Machines are better at fast arithmetic, sustained monitoring and repetition; people are better at judgement, improvisation and context. The form goes back to a 1951 report on air navigation and traffic control edited by Paul Fitts, and it has been reinvented, in identical shape, in every wave of automation since. It is intuitive, it is teachable, and it does not work.

Sidney Dekker and David Woods set out why in 2002, in a paper whose title asks whether the method is engineering or magic. Their argument is that quantitative "who does what" allocation cannot deliver coordination "because the real effects of automation are qualitative: it transforms human practice and forces people to adapt their skills and routines". They name the assumption underneath the lists the substitution myth: the idea that technology can be introduced "as a simple substitution of machines for people, preserving the basic system while improving it on some output measures".

Two sentences from that paper do more work than anything else in this literature. The first: "Capitalizing on some strength of automation does not replace a human weakness. It creates new human strengths and weaknesses, often in unanticipated ways." The second is the one that names the hole this page is trying to fill. Writing about the supervisory control levels that the field uses instead of Fitts lists, they observe that "neither the list nor much of the accompanying supervisory control literature explains the cognitive work that might be involved in deciding how and when to intervene or how to switch from level to level". That was written in 2002 about aircraft. It is a precise description of what is missing from every agent deployment plan in 2026.

Their replacement for the allocation question is a single line: "The question for successful automation is not 'who has control over what or how much'. It is 'how do we get along together'." Note that the Dekker and Woods paper is an argument rather than a study. It contains no data, no participants and no experiment, and it should be read as the field's best articulation of a problem rather than as evidence about its size.

Medicine has a name for the thing between the steps#

The discipline that has taken boundaries most seriously is patient safety, and its vocabulary transfers almost without translation. Cook, Render and Woods, writing in the BMJ in 2000, define the unit: "Gaps are discontinuities in care. They may appear as losses of information or momentum or interruptions in delivery of care."

Three of their observations belong in any agent design review. The first is that gaps are usually invisible because they are usually bridged: "most gaps are anticipated, identified, and bridged and their consequences nullified by the technical work done at the sharp end. These gap driven activities are so intimately woven into the fabric of technical work that neither outsiders nor insiders recognise them as distinct from other technical work." The second is that bridging is not solving: "To bridge a gap is not to eliminate it; some bridges are robust and reliable but others are frail, brittle, and easily undone by outside circumstances." The third is the one that inverts the standard safety argument: "accidents occur because conditions overwhelm or nullify the mechanisms practitioners normally use to detect and bridge gaps", and therefore "efforts to forestall errors by isolating practitioners from the system will misfire".

The example they choose to illustrate it was written about nursing in 2000 and could have been written about agents this year. Splitting nursing work between nurses and less credentialed patient care technicians has substantial economic benefit, because it lets the nurse concentrate on the tasks that require the credential. Among the side effects, they write, "are restrictions on the ability of the individual nurse to anticipate and detect gaps in the care of the patients. The nurse now has more patients to track, requiring more (and more complicated) inferences about which patient will next require attention, where monitoring needs to be more intensive, and so forth."

Read that with an agent in the place of the technician. The delegation is real, the saving is real, and the cost is a supervisor whose attention is now spread across more parallel streams and who has less of the direct contact from which anticipation was built. That is the same structure as the invisible work of oversight, arriving from clinical medicine twenty-five years early.

What happened when nine hospitals redesigned one boundary#

Handoffs are the one boundary anybody has run a serious intervention on. Starmer and colleagues implemented a structured handoff bundle across nine paediatric residency programmes in the United States and Canada between January 2011 and May 2013, with 875 consenting residents and 10,740 patient admissions across matched six-month periods.

Their headline: "the medical-error rate decreased by 23% from the preintervention period to the postintervention period (24.5 vs. 18.8 per 100 admissions, P<0.001), and the rate of preventable adverse events decreased by 30% (4.7 vs. 3.3 events per 100 admissions, P<0.001)." Near misses and non-harmful errors fell 21 per cent.

Four details from the same paper decide how much weight the result can carry, and they are the reason it is worth citing rather than quoting. Non-preventable adverse events did not move, at 3.0 against 2.8 per 100 admissions with P=0.79, which is what you would expect if the intervention was doing what it claimed and is the strongest internal evidence that it was. The improvement cost no time: oral handoff duration went from 2.4 to 2.5 minutes per patient, P=0.55, with no change in resident workflow. Error types split, with diagnostic and history-related errors falling significantly while medication, procedure, fall and infection errors did not. And error rates did not change significantly at three of the nine sites, even though written and oral handoff processes improved at all nine. The authors state plainly that the design "precludes definitively establishing a causal link", and that bundling the intervention "prevents us from determining which elements were most essential".

The transferable finding sits underneath the 23 per cent: a handoff can be made substantially safer without being made longer, by changing what is transferred rather than how much time is spent transferring it. The same change then fails to reproduce in a third of settings for reasons nobody has established.

One number that belongs to this territory is absent on purpose. The claim that 80 per cent of serious medical errors involve miscommunication during handoff is quoted constantly and is not in the Joint Commission's Sentinel Event Alert on inadequate hand-off communication, which was read in full for this page. The alert's quantified claims are that communication failures were responsible at least in part for 30 per cent of malpractice claims over five years, and that 69 per cent of clinical learning environments had no standardised handoff process. Both are footnoted to documents not read here, and both concern communication generally rather than handoff specifically.

Asiana 214, and the mode nobody read#

The clearest published account of a boundary failure is an accident report. On 6 July 2013 a Boeing 777 struck the seawall short of runway 28L at San Francisco. The NTSB found that the pilot flying selected a mode that produced a climb rather than the descent he wanted, then disconnected the autopilot and moved the thrust levers to idle. That input caused the autothrottle to change to HOLD, a mode in which it does not control airspeed.

The board's sentence is the whole argument of this page in eighteen words: "Neither the PF, the pilot monitoring (PM), nor the observer noted the change in A/T mode to HOLD." Three trained professionals, in a cockpit, on a clear day, did not register a state transition that had just handed them responsibility for something the machine had been doing. The NTSB's probable cause names, among the contributing factors, "the complexities of the autothrottle and autopilot flight director systems that were inadequately described in Boeing's documentation and Asiana's pilot training, which increased the likelihood of mode error", and attributes the insufficient monitoring in part to "automation reliance".

Nothing failed. The autothrottle entered HOLD correctly. The system did exactly what its logic specified. What was defective was the transfer: a boundary was crossed, authority moved from the machine to the people, and the people were not told in a way that reached them. This is one accident, N of one, and it cannot establish how often mode confusion occurs. What it can do is show what the failure looks like when the allocation was correct and the join was not.

The regulation names the capability and skips the moment#

European law has more to say about this than anything else on the statute book, and it still stops short. Article 14 of Regulation (EU) 2024/1689 requires that high-risk systems be designed so that the natural persons assigned to oversight are enabled to understand the system's capacities and limitations and monitor its operation, "to remain aware of the possible tendency of automatically relying or over-relying on the output produced by a high-risk AI system (automation bias)", to interpret the output correctly, "to decide, in any particular situation, not to use the high-risk AI system or to otherwise disregard, override or reverse the output", and "to intervene in the operation of the high-risk AI system or interrupt the system through a 'stop' button or a similar procedure that allows the system to come to a halt in a safe state".

That is a serious provision and it is genuinely unusual for naming automation bias in legislation. Read it against Asiana, though, and the gap is exact. Article 14 contains no occurrence of handoff, handover, transition or transfer of control. It is written entirely in the vocabulary of standing capability: the person must be able to intervene. It does not ask what the person knows at the instant they take over, whether the system announced that it had stopped doing something, or whether the safe state it comes to rest in is legible to whoever now owns it. A halt is not a handoff. The estate has argued the authority half of this in who can override an AI system and the design half in the stop button page; the moment of transfer is the piece neither the law nor those pages reach.

The only measurement of where multi-step systems break#

There is now one piece of work with real numbers on where agentic processes fail. Cemri and colleagues annotated more than 1,600 execution traces across seven multi-agent frameworks, building the taxonomy from 150 traces with expert annotators at an inter-annotator kappa of 0.88. Their failure distribution runs: system design issues 41.8 per cent, inter-agent misalignment 36.9 per cent, task verification 21.3 per cent.

Inter-agent misalignment is the boundary category. They define it as failures that "arise from a breakdown in critical information flow from inter-agent interaction and coordination during execution", and break it down into unexpected conversation resets at 2.20 per cent, proceeding with wrong assumptions rather than seeking clarification at 6.80 per cent, task derailment at 7.40 per cent, withholding crucial information at 0.85 per cent, ignoring another agent's input at 1.90 per cent, and mismatches between reasoning and action at 13.2 per cent. More than a third of observed failure sits at the joins, and the largest single mode is a system doing something other than what it had just reasoned.

Their second finding matters more for anyone buying a solution. Protocols do not fix it: "the errors we observe in FC2 occur even when agents within the same framework communicate using natural language", and their stated insight is that "solutions focused on context or communication protocols are often insufficient for FC2 failures". Standardising the message format does not make the sender say the useful thing.

Now the limitation, which is also the finding. Every handoff in that dataset is agent to agent. There are no humans in it and no human checkpoints anywhere. A search for 2024 to 2026 work measuring long-horizon agent performance with a person as a step in the process returned nothing usable. So the field has begun to quantify where machine-to-machine boundaries fail and has not started on the boundary this page is about, which is the one every accountability model in the world depends on.

Four questions to put at every boundary#

The estate's delegation boundary map answers the task question: nine stages, four tests, and where a given piece of work should sit. These four questions sit on top of it and are asked once per join rather than once per task, which is the change of unit this page is arguing for. They are assembled from Cook, Render and Woods on gaps, from Starmer and colleagues on what a structured transfer actually contains, and from what Article 14 leaves out.

One consequence of asking these questions is unwelcome and should be said. Fewer boundaries beat better boundaries. Every join is a place to lose something, and a process that hands work back and forth six times has six opportunities to lose it, however well each handoff is designed. Where the choice exists, the end-to-end answer is usually to consolidate the crossings rather than to instrument all of them.

Where this argument came from#

In Agentic AI, published on 8 September 2024, Rahim Hirji described the shift this page is about before the word had settled: agentic systems are "the difference between an intern who waits for instructions and a colleague who sees what needs to be done and does it". He then put the management question that the field is still avoiding: "If AI becomes more like a colleague than a tool, how do we manage it? Do we have to train and guide AI agents the way we would new employees?" And, in the same piece, "it's not far-fetched to think AI could even belong on an org chart."

Eleven months later, in Rules Before Tools on 17 August 2025, the second of ten rules names the design object directly. "Pouring AI into yesterday's process just scales yesterday's problems," and the instruction attached to it is to sketch the current process, circle the delays, and "rebuild 1 flow with fewer handoffs and smarter checkpoints". That is a dated instance of treating handoffs, rather than tasks, as the thing to redesign, written a year before the multi-agent failure taxonomies arrived at the same place with data.

The concern underneath both is the one this research keeps returning to. In Invisible Work in 2025 he argued that the checking, the noticing and the quiet correction that keep an organisation upright have never appeared on any measure of output, which makes them the easiest to cut and the most expensive to lose. Cook, Render and Woods made the same argument about clinicians in 2000 and called it bridging gaps. It is the same work, and an end-to-end redesign that does not account for it will remove it without ever having seen it.

What could not be established here#

There is no measurement of human-agent handoff failure. Everything quantitative on this page is either clinical, aeronautical, or agent-to-agent, and each transfer to the AI case is an argument rather than a finding. A page claiming otherwise would be inventing a literature.

Three sources are cited here more thinly than this estate prefers. Saying so beats letting the citation imply a reading. Bainbridge's Ironies of Automation is graded in the evidence base and its argument is used, but the full text could not be opened for this page, so nothing is quoted from it. Parasuraman, Sheridan and Wickens' four-stage model is paywalled; the stage names and the levels of automation are quoted here as Dekker and Woods render them, not from the original. The BEA final report on Air France 447, the other obvious case in this territory, could not be reached at the investigating authority's own domain, so it is not used and Asiana carries the section alone.

Finally, the four questions above have not been tested. They are derived from research on adjacent boundaries and from an accident report, and their status is a design proposal. The I-PASS result is the nearest thing to evidence that structuring a transfer changes outcomes, and its subject is people handing over to people.

Key sources

On dividing the tasks rather than the process, the delegation boundary map, which tasks workers do not want automated and what stays human. On the people at the boundary, who manages AI agents, who supervises work they cannot do and the invisible work of oversight. On the moment of intervention, how to design a stop button people will use, when to override AI and why human in the loop is not a safeguard. On the shape of the process itself, what happens to work that moves information and the shape of the organisation after AI.

About this research#

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. The substitution myth is Dekker and Woods' term. Gaps and bridging are Cook, Render and Woods'. The failure categories and their percentages are Cemri and colleagues'. Function allocation, supervisory control and mode error are established vocabulary and belong to nobody. Nothing on this page is a SuperSkills coinage. The Fitts list is credited to Fitts as editor of the 1951 report rather than as its sole author, which is how the National Research Council catalogued it. Article 14 was read at the European Commission's own AI Act Service Desk, because EUR-Lex returned an empty document to every route tried during this build; the Service Desk states its text is the official version of 13 June 2024. One widely repeated figure, that 80 per cent of serious medical errors involve handoff miscommunication, was searched for in the Joint Commission alert it is usually attributed to and is not there, so it does not appear above.

How this research works  ·  Reviewed quarterly  ·  Found an error? Tell me and it is corrected on the page.

Evidence review · SS-2026-169 · Graded against the published rubric

Cite this page

Hirji, R. (2026). How do humans and agents divide work across a process?. The SuperSkills evidence base, SS-2026-169. https://thesuperskills.com/research/how-do-humans-and-agents-divide-work. Last reviewed 4 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

How do humans and agents divide work across a process?

Not by dividing it. The evidence from fifty years of function allocation is that assigning tasks to whichever party is better at them does not produce a working system, because each assignment creates new work for the other party that did not exist before. Dekker and Woods call this the substitution myth. The design object is the boundary rather than the task: at each point where work passes between a person and a machine, what state is transferred, whether the transfer is announced, whether the receiver can reconstruct the reasoning, and who is accountable on either side. Those questions have answers in clinical handover research and in accident investigation, and almost none in the AI literature.

What is the substitution myth?

The assumption that new technology can be introduced as a straight swap of machines for people, leaving the rest of the system intact but improved. Dekker and Woods argue it is false: 'Capitalizing on some strength of automation does not replace a human weakness. It creates new human strengths and weaknesses, often in unanticipated ways.' They add that allocating a function creates new functions for the other partner that did not exist before, such as searching for the right display page. The term is theirs and is not a SuperSkills coinage.

Where do multi-step agent processes actually fail?

In the best-documented dataset, at the joins. Cemri and colleagues annotated more than 1,600 traces across seven multi-agent frameworks and classify failures as system design issues 41.8 per cent, inter-agent misalignment 36.9 per cent and task verification 21.3 per cent. Inter-agent misalignment is defined as a breakdown in critical information flow during execution, and its largest single component is a mismatch between an agent's reasoning and its action, at 13.2 per cent. The authors also report that communication protocols do not fix it, because the errors occur even when agents in the same framework communicate in natural language. Every handoff studied is agent to agent: no humans appear anywhere in that work.

Does the EU AI Act require anything about handoffs between AI and humans?

No. Article 14 of Regulation (EU) 2024/1689 requires that human overseers be enabled to monitor the system, remain aware of automation bias, interpret the output, decide to disregard or override it, and interrupt the system through a stop button that brings it to a halt in a safe state. It says nothing about handoff, handover, transition or transfer of control. It mandates that a person be able to intervene, and is silent on what state that person inherits at the moment they do.

In this hub

Organisations and leadership

What a leadership team actually has to decide, and what to measure.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Task-level allocation has failed since 1951. Every assignment creates new work for the other side, and the gaps are where the failures live. Settling the boundary, and who owns those gaps, is what I come in to do. AI advisory for CEOs and boards.

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.