Here is the finding that should reorganise how organisations think about oversight, and which almost nobody in the field has absorbed. Across 106 experimental studies and 370 effect sizes, putting a human and an AI together produced decisions that were on average worse than the better of the two working alone. Not worse than the human. Worse than whichever of the two was better at that task. The losses concentrated specifically in decision-making, while content creation showed real gains. Keep a human in the loop is the answer everyone gives to this, and the evidence does not support it. What survives is narrower: pairing helps when the human is better than the machine at the task, and hurts when the machine is better and the human overrides it or fails to catch it. Which means the whole discipline lives upstream, in working out which of you is actually better at what, before the decision arrives.
What the meta-analysis found#
The anchor result is a preregistered systematic review and meta-analysis by Vaccaro, Almaatouq and Malone, published in Nature Human Behaviour in 2024. They gathered 106 experimental studies reporting 370 effect sizes and asked a simple question: does a human working with an AI outperform the best of a human alone or an AI alone? On average, no. The combination performed significantly worse, with an effect size of Hedges' g of -0.23. The headline is arresting, but the structure underneath it is what matters. Losses were concentrated in decision-making tasks and gains were significantly larger in content-creation tasks. And the direction was predictable: where humans alone outperformed the AI, the combination gained; where the AI alone outperformed humans, the combination lost. Pairing did not average the two. It frequently dragged the better performer down towards the worse one.
That mechanism has a long-established name. Parasuraman and Manzey, reviewing decades of work across aviation, medicine and the military in 2010, described automation bias and complacency: the tendency to under-question automated advice, present in novices and experts alike, resistant to training, and worse under time pressure and high workload. Ignorance has nothing to do with it. This is what attention does when a competent system is doing the monitoring for you.
The most quoted demonstration in a professional setting is the jagged technological frontier. In 2023, Dell'Acqua and colleagues, working with Boston Consulting Group and researchers at Harvard, MIT and Wharton, gave 758 consultants access to GPT-4. Inside the model's competence, AI-assisted consultants were dramatically better and faster. On a task deliberately designed to sit just outside it, consultants using AI performed worse than consultants with no AI at all. The frontier is jagged rather than smooth, which is the difficult part: you cannot infer from a model's brilliance on one task that it is competent on an adjacent one, and the confidence of its output tells you nothing about which side of the line you are on.
Two findings from the trust literature explain why calibration is so hard, and they appear to contradict each other until you look at when each applies. Dietvorst, Simmons and Massey, writing in the Journal of Experimental Psychology: General in 2015, described algorithm aversion: after seeing an algorithm err, people abandon it, even when it demonstrably outperforms them and even when they have just watched their own worse performance. Logg, Minson and Moore, in 2019, described the opposite tendency, algorithm appreciation: for many estimates and forecasts, people weight algorithmic advice more heavily than human advice, with experts in the domain being the notable exception. Put together, they show trust moving for reasons unrelated to accuracy. Too much trust or too little turns out to be the wrong axis. And a policy instructing people to use judgement cannot fix a miscalibration it never diagnoses.
Finally, the gains are real and they are unevenly distributed. Brynjolfsson, Li and Raymond, studying 5,172 customer-support agents, found AI assistance raised productivity by fifteen percent on average, thirty percent for the newest and least experienced staff, and almost nothing for the most skilled. And the 2025 Microsoft Research and Carnegie Mellon survey of 319 knowledge workers found that the higher a worker's confidence in the AI, the less critical thinking they reported applying, with effort shifting from producing the work to verifying the output.
The studies stop at June 2023#
The meta-analysis covers studies published between January 2020 and June 2023, which means most of the underlying work predates the current generation of frontier models. If model capability has risen since, the balance of "who is better at this task" has shifted in the machine's favour, and the practical implication changes: more tasks fall into the category where human intervention subtracts rather than adds. That does not weaken the finding. It sharpens the warning, because it means more of the decisions where a human is nominally in the loop are decisions where the human is nominally in the way.
There is also a fairness point about the benchmark. "Worse than the best of either alone" is a demanding comparison, because in real settings you rarely know in advance which of the two is better at the specific task in front of you. That uncertainty is why the finding matters, though a combination which underperforms an oracle-selected best performer may still beat the realistic alternative of picking one and hoping. So the reading to take away is narrower than "human-AI teams are useless". Pairing has to be designed against a known division of competence. Undesigned, it comes out worse than either party alone.
Most of the underlying studies are laboratory or bounded-task experiments, with short horizons and clear right answers. Real organisational decisions are longer, more ambiguous, more political, and rarely scored. And none of this measures the second-order effect: what happens to a person's decision-making capability after two years of supervising a machine rather than deciding. That question is taken up in how humans learn with AI.
Human in the loop describes nothing#
Human in the loop is a reassurance rather than a design. It has become the phrase organisations reach for at the moment they would otherwise have to think. Putting a person at the end of an automated process, with no time budget, no authority to stop it, no stated basis on which they would disagree, and no consequence if they simply approve, does not produce oversight. It produces a signature. The meta-analysis is the empirical case for something I have argued for years from practice: the presence of a human does not improve a decision, and a human who has been positioned to rubber-stamp will make the system worse than either party alone, because they add latency and the appearance of scrutiny without the substance.
So the useful question is where the human sits. Reviewing at the end is the weakest position available: the framing is already set, the options have been narrowed, anchoring has happened, and the effort required to reopen the question is far higher than the effort required to approve. Being present at the start is a different job entirely, because that is where the problem is defined, the constraints are set, the success criteria are chosen and the alternatives that will never be generated are excluded. That is the principle I call Human at the Start, and the evidence on framing effects and automation bias makes it more than a stylistic preference. The start is the only point in the process where a human has leverage that is cheaper to exercise than to skip.
Read the Vaccaro result carefully and it yields three conditions under which pairing genuinely helps, which is more useful than the headline. Pairing helps when the human is better at the task, so the machine augments rather than leads. It helps when the task is generative rather than evaluative, because content creation showed the gains while decision-making showed the losses. And it helps when the human's contribution is defined as something other than approval, such as setting the problem, supplying context the model cannot have, or making the value trade-off the model has no standing to make. Where none of those three holds, you are not designing a partnership. You are adding a person to a process for comfort, and paying for it in accuracy.
The commercial consequence follows directly. If verification is the human's job, then verification is the skilled work and should be resourced, trained and paid as such. Almost no organisation does this, because checking looks like an administrative act and producing looks like a professional one. I call that mispricing the verifier's discount, and the jagged-frontier result is what it costs.
Deciding who decides: a working allocation#
Before a class of decision reaches a workflow, answer four questions in writing. The point of writing them down is not bureaucracy; it is that an organisation which has never written them has already answered them by default.
- Who is better at this, honestly? Not who should be. On this specific task, with this data, is the model more accurate than your people, less accurate, or unknown? Unknown is a legitimate answer and it changes the design, because it means you are running an experiment and should measure it as one.
- What is the human actually adding? Name it: framing, context, the value judgement, accountability, or catching a specific known failure mode. If the answer is "confidence", remove them from the loop and put them at the start instead.
- What would make a person disagree, and can they? Specify in advance the conditions under which the output should be rejected, and confirm the person has the time, the standing and the authority to reject it. Oversight without stop-work authority is theatre.
- How would we know this is going wrong? Approval rates near a hundred percent are a signal, not a success. Sample decisions and score them independently, because a process where nobody ever disagrees is indistinguishable from a process where nobody is looking.
For decisions taken by agents rather than assistants, where the system plans and acts rather than advises, the allocation question becomes sharper still and is developed separately in AI agents and human judgement.
What this looks like in practice#
A credit team runs a model that is measurably better than its analysts at scoring standard applications, and worse at the unusual ones. The wrong design is a human reviewing every case, which adds delay and drags the good decisions towards the human's lower accuracy. The right design routes the standard cases to the model with sampled audit, and routes the unusual ones to a person before the model has framed them. Same technology, opposite performance, and the difference is one page of thinking nobody had time for.
A clinical team uses an AI triage tool. The failure to design against is not the model being wrong; it is the model being right often enough that questioning it starts to feel like obstruction. That is the automation-complacency finding, and the countermeasure is structural rather than attitudinal: a stated disagreement rate that is expected to be non-zero, protected time to examine flagged cases, and no professional cost for overriding.
A leadership team receives an AI-generated market analysis, discusses it for forty minutes, and approves the recommendation. Nobody asks what question was put to the model, what it was not given, or which options it never generated. The meeting felt rigorous. Every input to the decision was set before anyone in the room was involved.
Say who, where, and with what authority#
Stop saying human in the loop and start saying who, where and with what authority. The phrase has become a way of not deciding. Replace it in your policies with a named person, a named point in the process, and a stated basis for disagreement.
Move the human upstream. Framing, constraints and success criteria are where a person still has leverage. Review at the end is where they have the least. It is where almost every organisation has put them.
Establish which of you is better, task by task, and write it down. This is the single highest-return piece of work available and it is nearly always skipped, because it requires admitting that on some tasks the machine is better and on others your people are, and both admissions are politically awkward.
Resource verification like the skilled work it is. Give it time, training, seniority and status, or accept that you have bought a process which fails exactly when it matters, on the tasks that sit outside the model's competence.
Measure disagreement. Track how often humans override, and investigate when the rate approaches zero. A hundred percent approval evidences an unmeasured process rather than a good model.
Development of the idea#
I argued in Entrepreneur UK in July 2026 that AI does not create bad decisions, it exposes them faster, which is the same point from the organisational side: AI reveals whether a decision process was ever a process. The accountability form of the argument, including the distinction between Human at the Start, human in the loop and human at the end, is set out in the European Business Review piece on accountability gaps in leadership decisions (21 August 2026). In the Observer in July 2026 I argued that judgement, not model-building, would define the next power class. The argument is set out at length in SuperSkills (Kogan Page, 2026).
What regulators have already decided about oversight#
The argument on this page has a legal counterpart, and it has been settled more clearly in European regulation than in most corporate policy. The test the courts apply is competence and an actual decision, not presence.
In 2023 the Amsterdam Court of Appeal ruled against Uber in a case brought by drivers deactivated after fraud flags. What the court examined was not the presence of a review step but whether the reviewer's qualifications and knowledge could be evidenced. Uber could not say who had decided or what they knew. In 2024 Italy's data protection authority, the Garante, fined Foodinho, the Glovo subsidiary, five million euros, a second offence after a 2.6 million euro fine in 2021. The remedy is the interesting part: reviewers must be adeguatamente formati, adequately trained. Competence, not attendance, became the standard.
In August 2026 the Dutch data protection authority fined Uber roughly 825 million euros over driver deactivations decided without human assessment, a decision Uber is appealing. The finding rests on an absence rather than an act: scoring drivers algorithmically remained lawful, and deciding their livelihood without a person did not.
There is a lesson here for anyone writing an AI policy. "A human reviews the output" is not a control that survives contact with a regulator, and increasingly not one that survives contact with a court. Who, with what training, deciding what, on what basis, with what authority to refuse: those are the questions being asked, and an organisation that cannot answer them has documented oversight rather than exercised it.
Nobody can tell you who benefits#
There is a further complication for anyone designing an oversight policy. It is the most awkward finding in this whole literature. Yu and colleagues, publishing in Nature Medicine in 2024, gave 140 radiologists AI assistance across fifteen chest X-ray tasks, roughly 5,190 observations, randomised, with statistical treatment designed to separate genuine individual differences from noise.
The effects diverged sharply. AI assistance helped some radiologists substantially and made others measurably worse. And nothing predicted which: not years of experience, not subspecialty, not prior familiarity with AI. Lower performers did not reliably gain, which is the assumption most deployment plans rest on.
The implication is uncomfortable, so let me put it plainly. Rolling a tool out to everyone in a professional group will help some of them and harm others, and at present there is no way to know in advance who is in which group. That is an argument for measuring individual effects rather than assuming an average. It is a further reason why "we have kept a human in the loop" describes a hope rather than a control.
What the automation literature settled decades ago#
Almost every finding above has a precedent in human-factors research, and reading it is the cheapest available upgrade to how an organisation designs oversight. Parasuraman and Riley, writing in Human Factors in 1997, set out a taxonomy that is still sharper than the vocabulary in current AI policy: use, misuse, disuse and abuse. Misuse is over-reliance. Disuse is unwarranted rejection, which is a real and separate failure. Abuse is deployment by designers in ways that ignore the human consequences. Collapsing all three into "over-reliance", as most AI guidance does, loses the distinctions that determine what you should actually do.
Skitka, Mosier and Burdick demonstrated automation bias experimentally in 1999 and split it into two kinds of error, and the split matters more than the phenomenon. Commission errors are acting on a wrong recommendation. Omission errors are missing something the system did not flag. Nearly every oversight process is designed to catch the first. The second is invisible by construction, because nothing appears on the screen to check. That is the failure mode that accumulates.
The most uncomfortable finding is Dzindolet and colleagues, in 2003. Trust mediates reliance, as you would expect. But they also found that explaining why an automated aid might err increased reliance on it, even when that restored trust was not warranted. That is a direct warning about explainability as a safety measure: telling people how a model can fail may make them trust it more rather than less. If your governance rests on transparency producing appropriate scepticism, this study says the mechanism can run backwards.
For deciding how far to automate rather than whether, Parasuraman, Sheridan and Wickens set out a four-stage, ten-level model in 2000 that remains more rigorous than most current thinking about what to delegate to an agent. And on the practical side, Daugherty and Wilson's Human + Machine names the hybrid roles in what they call the missing middle, where humans train, explain and sustain machine systems. The NIST AI Risk Management Framework gives this a governance vocabulary a board will recognise, though its weakness is instructive: it can be satisfied procedurally by an organisation that documents oversight without exercising it.
Key research and primary sources
- Garante per la protezione dei dati personali (2024). Provvedimento nei confronti di Foodinho s.r.l. (Glovo).
- Yu, F., Moehring, A., Banerjee, O., Salz, T., Agarwal, N. and Rajpurkar, P. (2024). Heterogeneity and predictors of the effects of AI assistance on radiologists. Nature Medicine, 30(3), 837-849. graded entry.
- Daugherty, P. R. and Wilson, H. J. (2024). Human + Machine: Reimagining Work in the Age of AI. Harvard Business Review Press, updated and expanded edition.
- Dzindolet, M. T., Peterson, S. A., Pomranky, R. A., Pierce, L. G. and Beck, H. P. (2003). The role of trust in automation reliance. International Journal of Human-Computer Studies, 58(6), 697-718.
- OECD (2025). OECD AI Capability Indicators: Technical Report. OECD Publishing, Paris, November 2025.
- Parasuraman, R. and Riley, V. (1997). Humans and Automation: Use, Misuse, Disuse, Abuse. Human Factors, 39(2), 230-253.
- Parasuraman, R., Sheridan, T. B. and Wickens, C. D. (2000). A Model for Types and Levels of Human Interaction with Automation. IEEE Transactions on Systems, Man and Cybernetics, Part A, 30(3), 286-297.
- Skitka, L. J., Mosier, K. L. and Burdick, M. (1999). Does automation bias decision-making?. International Journal of Human-Computer Studies, 51(5), 991-1006.
- Vaccaro, M., Almaatouq, A. and Malone, T. (2024). When combinations of humans and AI are useful: a systematic review and meta-analysis. Nature Human Behaviour, 8, 2293-2303.
- Dell'Acqua, F. et al. (2023). Navigating the Jagged Technological Frontier. Harvard Business School and BCG working paper.
- Parasuraman, R. and Manzey, D. H. (2010). Complacency and Bias in Human Use of Automation. Human Factors, 52(3).
- Dietvorst, B. J., Simmons, J. P. and Massey, C. (2015). Algorithm Aversion: People Erroneously Avoid Algorithms After Seeing Them Err. Journal of Experimental Psychology: General, 144(1).
- Logg, J. M., Minson, J. A. and Moore, D. A. (2019). Algorithm Appreciation: People Prefer Algorithmic to Human Judgment. Organizational Behavior and Human Decision Processes, 151, 90-103.
- Brynjolfsson, E., Li, D. and Raymond, L. (2025). Generative AI at Work. Quarterly Journal of Economics, 140(2), 889-942. Earlier version NBER Working Paper 31161.
- Lee, H.-P. et al. (2025). The Impact of Generative AI on Critical Thinking. Microsoft Research and Carnegie Mellon, CHI 2025.
Related SuperSkills research#
On where the human belongs in the process, Human at the Start and AI agents and human judgement. On the underlying capability, AI and human judgement and decision quality in the AI era. On the mispricing of verification, the verifier's discount. On the organisational conditions, drift versus design and AI workforce strategy. The leadership framing is how leaders should respond to AI. The graded evidence is in the evidence base. On the underlying tendency, automation bias; on the meta-analytic baseline, what is human-AI collaboration? The operational version, stage by stage with a downloadable grid, is the Delegation Boundary Map. On who carries the verification duty, who owns verification when AI does the work. The position that follows from this, put simply, is that human in the loop is not a safeguard.
About this research#
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. This work draws on research across more than 200 organisations in 30 countries over seven years. Findings are attributed to the studies that produced them and kept separate from the interpretation, which is the author's. Automation bias, automation complacency, algorithm aversion and algorithm appreciation are established concepts from the research literature and are not his. Human at the Start, drift versus design and the verifier's discount are part of the SuperSkills lexicon. This is a living reference, reviewed and updated as significant new evidence appears.
Cite this
Hirji, R. (2026). Human and AI decision making. The SuperSkills Intelligence Company. Last reviewed 26 August 2026. thesuperskills.com/research/human-ai-decision-making
