← Research
Research

What is human-AI collaboration?

On average it underperforms the stronger half. Where it works, and why the common configuration is the one the evidence says is worst.

Last reviewed: 26 August 2026

What is human-AI collaboration, and does it actually work? This page gives the definition, the meta-analytic baseline most deployments never test against, and the design change Rahim Hirji argues turns a losing configuration into a winning one.

Human-AI collaboration is any arrangement where a person and an automated system contribute to the same piece of work. The honest definition has to begin with an inconvenient finding: on average, it does not work. A 2024 meta-analysis in Nature Human Behaviour pooled 370 effect sizes from 106 experiments and found that human-AI combinations performed worse, on average, than the better of the human alone or the AI alone. Not worse than both. Worse than whichever was stronger. That is the baseline any claim about collaboration has to clear, and most organisational deployments never test against it.

Which does not mean collaboration is a myth. The same analysis found real synergy in a specific place, and the pattern of where it appears and where it disappears is the most useful thing anyone has published on this question.

Definition

Human-AI collaboration: a work arrangement in which a person and an automated system each contribute to a shared output, with the division of labour, the point of human entry, and the basis on which the human may override the system all specified in advance. Where those three things are unspecified, the arrangement is not collaboration but sequential handover, and the evidence suggests it usually underperforms whichever party was stronger.

What the evidence shows

Vaccaro, Almaatouq and Malone conducted the systematic review, covering studies published between 2020 and 2023. The headline average loss is striking, but the disaggregation matters more.

Synergy appeared reliably in creation tasks, where the human and the system are producing something. It failed in decision tasks, where the human is judging whether the system is right. And the strongest predictor of which way a study went was the baseline: when the human alone outperformed the AI alone, combining them helped; when the AI alone outperformed the human, combining them dragged the result down towards the human's level. In other words, human oversight of a system already better than you tends to subtract.

That is not an argument for removing oversight, and the medical evidence shows why. Yu and colleagues, in Nature Medicine in 2024, found the effect of AI assistance on radiologists varied enormously between individuals, was not predicted by experience or by prior AI exposure, and helped some readers while degrading others. An average effect concealed opposite effects on different people. Any policy of the form "clinicians will use the tool" is therefore a policy with unknown sign.

The mechanism behind decision-task failure is well documented and is not new. Bainbridge described the ironies of automation in 1983: automate the routine and you leave the human with the hardest residue, monitoring, while removing the practice that made them able to do it. Skitka and colleagues measured the two resulting error types directly in 1999, omission and commission. Parasuraman and Manzey's 2010 review established that these are properties of attention allocation rather than of carelessness. None of this required generative AI to be discovered, and generative AI has not repealed it.

Where the evidence is uncertain

The meta-analysis covers experiments up to 2023, using systems less capable than today's, and its tasks are mostly short and laboratory-based. Long-horizon collaboration, where a person and a system work together over weeks and the person learns the system's failure modes, is barely represented. It is plausible that experienced pairings behave quite differently from first-encounter ones, and equally plausible that they behave worse because familiarity increases complacency. Nobody has the data.

A fair objection also applies to the creation-versus-decision split: it is a pattern found across heterogeneous studies rather than a mechanism anyone has isolated experimentally. It should be treated as a strong working hypothesis, not a law.

The SuperSkills view

Most organisations are running the configuration the evidence says is worst, and calling it responsible practice. The machine drafts, the human reviews, and the human is described as being in the loop. That is a decision task, performed on output the reviewer did not generate, usually under time pressure, frequently by someone who could not have produced the work themselves. It is the exact shape the meta-analysis found to underperform, and it is the default because it is the easiest to install rather than because it works.

The alternative I argue for is Human at the Start: the human sets the question, the constraints and the rejection criteria before the system produces anything. This is not a preference for humans over machines. It converts a decision task, where combination subtracts, into a creation task, where the evidence says it adds. It also gives the reviewer a position to compare against, which is what makes review something other than a fluency check.

Two corollaries follow, and they are the ones organisations skip. First, if the system is genuinely better than the human at a task, the honest options are to let it run without theatre and accept the accountability, or to keep the human and accept that you are paying for a slower and less accurate result to preserve something else you value, such as legitimacy or the ability to explain a decision. Both are defensible. Pretending you are getting the best of both is not, and that pretence is what I call usage theatre.

Second, collaboration has a maintenance cost nobody budgets. The human half of the pairing has to stay capable enough to disagree, which means retaining practice at work the machine is doing. Left alone, that capability erodes silently, which is capability debt. A collaboration design that does not include how the human stays good at the thing is a collaboration design with an expiry date on it.

Designing a pairing that clears the baseline

Related SuperSkills research

On decision architecture, human and AI decision making. On the underlying tendency, automation bias and how to know when AI is wrong. On the organisational choice, drift versus design and AI workforce strategy. The claims themselves, banded by evidence strength and including what remains unknown, are in what we actually know about AI and human capability.

Key research and primary sources

About this research

Rahim Hirji is the author of SuperSkills: The Seven Human Skills for the Age of AI (Kogan Page, 2026) and the founder of The SuperSkills Intelligence Company. Human-AI collaboration is an established field term and is not his coinage; Human at the Start, usage theatre and capability debt are. Findings are attributed to the studies that produced them and kept separate from the interpretation. The graded evidence, including what each study does not support, is in the evidence base.

Cite this

Hirji, R. (2026). What is human-AI collaboration? The SuperSkills Intelligence Company. Last reviewed 26 August 2026. thesuperskills.com/research/what-is-human-ai-collaboration

In this hub

AI and Human Judgement

Does AI weaken judgement? The evidence, and what to do about it.

The work

Where the writing comes from.

These essays draw on research across more than 200 organisations in 30 countries. See the wider body of work, or bring it into your organisation.

All research →