← Research
Research

What is human-AI collaboration?

On average it underperforms the stronger half. Where it works, and why the common configuration is the one the evidence says is worst.

Last reviewed: 26 August 2026

What is human-AI collaboration, and does it actually work? This page gives the definition, the meta-analytic baseline most deployments never test against, and the design change Rahim Hirji argues turns a losing configuration into a winning one.

Question this page answersAll 616 questions this research covers

Human-AI collaboration is any arrangement where a person and an automated system contribute to the same piece of work. The honest definition has to begin with an inconvenient finding: on average, it does not work. A 2024 meta-analysis in Nature Human Behaviour pooled 370 effect sizes from 106 experiments and found that human-AI combinations performed worse, on average, than the better of the human alone or the AI alone. Not worse than both. Worse than whichever was stronger. That is the baseline any claim about collaboration has to clear, and most organisational deployments never test against it.

Which does not mean collaboration is a myth. The same analysis found real synergy in a specific place, and the pattern of where it appears and where it disappears is the most useful thing anyone has published on this question.

Definition#

Human-AI collaboration: a work arrangement in which a person and an automated system each contribute to a shared output, with the division of labour, the point of human entry, and the basis on which the human may override the system all specified in advance. Where those three things are unspecified, the arrangement is not collaboration but sequential handover, and the evidence suggests it usually underperforms whichever party was stronger.

What the systematic review covered#

Vaccaro, Almaatouq and Malone conducted the systematic review, covering studies published between 2020 and 2023. The headline average loss is striking, but the disaggregation matters more.

Synergy appeared reliably in creation tasks, where the human and the system are producing something. It failed in decision tasks, where the human is judging whether the system is right. And the strongest predictor of which way a study went was the baseline: when the human alone outperformed the AI alone, combining them helped; when the AI alone outperformed the human, combining them dragged the result down towards the human's level. In other words, human oversight of a system already better than you tends to subtract.

That is not an argument for removing oversight, and the medical evidence shows why. Yu and colleagues, in Nature Medicine in 2024, found the effect of AI assistance on radiologists varied enormously between individuals, was not predicted by experience or by prior AI exposure, and helped some readers while degrading others. An average effect concealed opposite effects on different people. Any policy of the form "clinicians will use the tool" is therefore a policy with unknown sign.

The mechanism behind decision-task failure is well documented and is not new. Bainbridge described the ironies of automation in 1983: automate the routine and you leave the human with the hardest residue, monitoring, while removing the practice that made them able to do it. Skitka and colleagues measured the two resulting error types directly in 1999, omission and commission. Parasuraman and Manzey's 2010 review established that these are properties of attention allocation rather than of carelessness. None of this required generative AI to be discovered, and generative AI has not repealed it.

The experiments stop at 2023#

The meta-analysis covers experiments up to 2023, using systems less capable than today's, and its tasks are mostly short and laboratory-based. Long-horizon collaboration, where a person and a system work together over weeks and the person learns the system's failure modes, is barely represented. It is plausible that experienced pairings behave quite differently from first-encounter ones, and equally plausible that they behave worse because familiarity increases complacency. Nobody has the data.

A fair objection also applies to the creation-versus-decision split: it is a pattern found across heterogeneous studies rather than a mechanism anyone has isolated experimentally. It should be treated as a strong working hypothesis, not a law.

The worst configuration, called responsible#

Most organisations are running the configuration the evidence says is worst, and calling it responsible practice. The machine drafts, the human reviews, and the human is described as being in the loop. That is a decision task, performed on output the reviewer did not generate, usually under time pressure, frequently by someone who could not have produced the work themselves. It is the exact shape the meta-analysis found to underperform. It is the default because it is the easiest to install rather than because it works.

The alternative I argue for is Human at the Start: the human sets the question, the constraints and the rejection criteria before the system produces anything. This is not a preference for humans over machines. It converts a decision task, where combination subtracts, into a creation task, where the evidence says it adds. It also gives the reviewer a position to compare against, which is what makes review something other than a fluency check.

Two corollaries follow, and they are the ones organisations skip. First, if the system is genuinely better than the human at a task, the honest options are to let it run without theatre and accept the accountability, or to keep the human and accept that you are paying for a slower and less accurate result to preserve something else you value, such as legitimacy or the ability to explain a decision. Both are defensible. Pretending you are getting the best of both is not, and that pretence is what I call usage theatre.

Second, collaboration has a maintenance cost nobody budgets. The human half of the pairing has to stay capable enough to disagree, which means retaining practice at work the machine is doing. Left alone, that capability erodes silently, which is capability debt. A collaboration design that does not include how the human stays good at the thing is a collaboration design with an expiry date on it.

Designing a pairing that clears the baseline#

On decision architecture, human and AI decision making. On the underlying tendency, automation bias and how to know when AI is wrong. On the organisational choice, drift versus design and AI workforce strategy. The claims themselves, banded by evidence strength and including what remains unknown, are in what we actually know about AI and human capability. The operational version, stage by stage with a downloadable grid, is the Delegation Boundary Map. On what oversight is now legally required to enable, meaningful human oversight. The position that follows from this, put simply, is that human in the loop is not a safeguard. On the definition, the jagged frontier.

Key research and primary sources

About this research#

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. Human-AI collaboration is an established field term and is not his coinage; Human at the Start, usage theatre and capability debt are. Findings are attributed to the studies that produced them and kept separate from the interpretation. The graded evidence, including what each study does not support, is in the evidence base.

How this research works  ·  Reviewed quarterly  ·  Found an error? Tell me and it is corrected on the page.

Cite this

Hirji, R. (2026). What is human-AI collaboration? The SuperSkills Intelligence Company. Last reviewed 26 August 2026. thesuperskills.com/research/what-is-human-ai-collaboration

Questions answered on this page

What is human-AI collaboration?

A work arrangement in which a person and an automated system each contribute to a shared output, with the division of labour, the point of human entry, and the grounds for human override all specified in advance. Where those three things are unspecified, the arrangement is sequential handover rather than collaboration. That distinction matters because the evidence shows unspecified handover usually underperforms whichever party was stronger on its own.

Does human-AI collaboration actually improve results?

Not on average. Vaccaro, Almaatouq and Malone pooled 370 effect sizes from 106 experiments for a 2024 Nature Human Behaviour meta-analysis and found human-AI combinations performed worse than the better of human alone or AI alone. Synergy did appear reliably in creation tasks, where the pair produces something together. It failed in decision tasks, where the human judges whether the system is right. The strongest predictor was the baseline: combining helped when the human alone outperformed the AI, and hurt when the AI alone outperformed the human.

Why does human in the loop review often fail?

Because it is a decision task performed on output the reviewer did not generate, usually under time pressure, and often by someone who could not have produced the work themselves. That is the configuration the meta-analysis found to underperform. The underlying mechanism was described by Bainbridge in 1983: automating the routine leaves the human with monitoring, the hardest residue, while removing the practice that made them capable of it.

How do you design human-AI collaboration that works?

Move the human to the front. Set the question, the constraints and the rejection criteria before the system generates anything, which converts a decision task into a creation task. Then test the pairing against the better of human alone and system alone rather than against nothing, state the override rule in advance, measure effects at the level of the individual rather than the average, and budget for keeping the human capable at the work they are supervising.

In this hub

Definitions

The terms this field uses, defined against their primary sources.

The work

Where the writing comes from.

These essays draw on research across more than 200 organisations in 30 countries. See the wider body of work, or bring it into your organisation.

All research →
Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory and coaching  ·  Enquire