A meta-analysis of 106 experiments found that human-AI combinations performed on average worse than the better of human or AI alone. The synergy premise underneath most adoption strategy is not supported by the largest evidence base available, and the exceptions are specific enough to design for.
The answer, in one line
On average, no. Vaccaro, Almaatouq and Malone's 2024 meta-analysis of 106 experiments in Nature Human Behaviour found human-AI combinations performed worse than the better of the two alone.
The finding#
Vaccaro, Almaatouq and Malone published a meta-analysis in Nature Human Behaviour in 2024 covering 106 experiments on human-AI teams. On average the combination underperformed the stronger of its two parts. Where the AI alone was better, adding a human made it worse. Where the human alone was better, the combination sometimes helped.
The pattern by task type is the useful part. Gains appeared in creation tasks, where the human is generating something and the model is supplying material. Losses concentrated in decision tasks, where somebody has to choose. That is close to the opposite of how most organisations have deployed these systems.
Why the combination fails#
The mechanism is described at appropriate reliance. To gain from the pairing, the human has to accept the machine's answer when it is right and override it when it is wrong, which requires being able to tell which is which, which usually requires the capability the tool was brought in to supply. Two decades of work have not produced an intervention that reliably delivers it.
The intuitive fix makes it worse. Bansal and colleagues found in 2021 that explanations increased acceptance of the model's answer without improving team accuracy, and acceptance rose for correct and incorrect outputs alike. Confidence scores move the problem rather than solving it, since models are frequently confident in the wrong places.
When it does help#
Four conditions recur across the literature, and none of them arrives by default.
- The human can verify. Checking has to be faster than producing, and possible at all. This is the first question of the use-or-keep test and the precondition for everything else.
- The human holds context the model does not. Where the distinguishing feature of the case is something you know and the model has never seen, the pairing has something to combine. Where the model has seen everything you have, it does not.
- The division is decided in advance. A boundary negotiated case by case under time pressure resolves towards accepting the output. Deciding the ground beforehand is what the three terrains are for.
- The human is better at the task alone. The meta-analysis found gains concentrated here, which inverts the usual deployment logic of putting AI where people struggle.
The field evidence is consistent with this. In the Brynjolfsson study of 5,172 support agents, productivity rose about 15 per cent on average, roughly 30 per cent among novices and almost nothing among the most experienced, and the workers followed only about 38 per cent of the AI's recommendations. The gain came from people exercising judgement about which suggestions to take, not from deference.
What the centaur story gets wrong#
The claim that human-machine teams beat both, borrowed from freestyle chess, is doing work it cannot support. Chess engines now dominate the pairing outright, and the meta-analytic evidence in knowledge work points the other way. The defensible position is that the combination is a design problem with identifiable conditions rather than a general property of putting the two together.
What follows for an organisation#
Sort tasks by whether the human or the machine is better alone, and be willing to act on the answer in both directions. Where the model is reliably better, the useful human role is interpretation, oversight and accountability rather than participation in the judgement, which is the first of the three terrains. Where the human is better, the model supplies material and the human decides. Where neither is clearly better, the combination is least likely to help and most likely to be adopted, because that is where it feels most useful.
What this does not establish#
The meta-analysis covers experiments, most of them short, many with non-expert participants and constructed tasks, which is a setting that can understate what a trained professional does with a tool over months. It also predates the current generation of systems. It does not show that human-AI collaboration cannot work, and it does show that it does not work by default, which is the assumption most deployment rests on.
Key sources
- Vaccaro, M., Almaatouq, A. and Malone, T. (2024). When combinations of humans and AI are useful. Nature Human Behaviour, 8.
- Bansal, G. et al. (2021). Does the whole exceed its parts? The effect of AI explanations on complementary team performance. CHI 2021.
- Brynjolfsson, E., Li, D. and Raymond, L. R. (2023). Generative AI at Work.
Related SuperSkills research#
Essay · SS-2026-347
Hirji, R. (2026). When does AI improve a decision?. The SuperSkills evidence base, SS-2026-347. https://thesuperskills.com/research/when-does-ai-improve-a-decision. Last reviewed 26 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work