Here is the finding that should reorganise how organisations think about oversight, and which almost nobody in the field has absorbed. Across 106 experimental studies and 370 effect sizes, putting a human and an AI together produced decisions that were on average worse than the better of the two working alone. Not worse than the human. Worse than whichever of the two was better at that task. The losses concentrated specifically in decision-making, while content creation showed real gains. So the honest answer to how humans and AI should decide together is not "keep a human in the loop", which is the answer everyone gives and which the evidence does not support. It is narrower: pairing helps when the human is better than the machine at the task, and hurts when the machine is better and the human overrides it or fails to catch it. Which means the whole discipline lives upstream, in working out which of you is actually better at what, before the decision arrives.
What the evidence shows
The anchor result is a preregistered systematic review and meta-analysis by Vaccaro, Almaatouq and Malone, published in Nature Human Behaviour in 2024. They gathered 106 experimental studies reporting 370 effect sizes and asked a simple question: does a human working with an AI outperform the best of a human alone or an AI alone? On average, no. The combination performed significantly worse, with an effect size of Hedges' g of -0.23. The headline is arresting, but the structure underneath it is what matters. Losses were concentrated in decision-making tasks and gains were significantly larger in content-creation tasks. And the direction was predictable: where humans alone outperformed the AI, the combination gained; where the AI alone outperformed humans, the combination lost. Pairing did not average the two. It frequently dragged the better performer down towards the worse one.
That mechanism has a long-established name. Parasuraman and Manzey, reviewing decades of work across aviation, medicine and the military in 2010, described automation bias and complacency: the tendency to under-question automated advice, present in novices and experts alike, resistant to training, and worse under time pressure and high workload. It is not ignorance. It is what attention does when a competent system is doing the monitoring for you.
The most quoted demonstration in a professional setting is the jagged technological frontier. In 2023, Dell'Acqua and colleagues, working with Boston Consulting Group and researchers at Harvard, MIT and Wharton, gave 758 consultants access to GPT-4. Inside the model's competence, AI-assisted consultants were dramatically better and faster. On a task deliberately designed to sit just outside it, consultants using AI performed worse than consultants with no AI at all. The frontier is jagged rather than smooth, which is the difficult part: you cannot infer from a model's brilliance on one task that it is competent on an adjacent one, and the confidence of its output tells you nothing about which side of the line you are on.
Two findings from the trust literature explain why calibration is so hard, and they appear to contradict each other until you look at when each applies. Dietvorst, Simmons and Massey, writing in the Journal of Experimental Psychology: General in 2015, described algorithm aversion: after seeing an algorithm err, people abandon it, even when it demonstrably outperforms them and even when they have just watched their own worse performance. Logg, Minson and Moore, in 2019, described the opposite tendency, algorithm appreciation: for many estimates and forecasts, people weight algorithmic advice more heavily than human advice, with experts in the domain being the notable exception. Put together, the picture is not that people trust machines too much or too little. It is that trust moves for reasons unrelated to accuracy, which is precisely the failure mode you cannot fix with a policy telling people to use judgement.
Finally, the gains are real and they are unevenly distributed. Brynjolfsson, Li and Raymond, studying 5,179 customer-support agents, found AI assistance raised productivity by fourteen percent on average, thirty-four percent for the newest and least experienced staff, and almost nothing for the most skilled. And the 2025 Microsoft Research and Carnegie Mellon survey of 319 knowledge workers found that the higher a worker's confidence in the AI, the less critical thinking they reported applying, with effort shifting from producing the work to verifying the output.
Where the evidence is uncertain
The meta-analysis covers studies published between January 2020 and June 2023, which means most of the underlying work predates the current generation of frontier models. If model capability has risen since, the balance of "who is better at this task" has shifted in the machine's favour, and the practical implication changes: more tasks fall into the category where human intervention subtracts rather than adds. That does not weaken the finding. It sharpens the warning, because it means more of the decisions where a human is nominally in the loop are decisions where the human is nominally in the way.
There is also a fairness point about the benchmark. "Worse than the best of either alone" is a demanding comparison, because in real settings you rarely know in advance which of the two is better at the specific task in front of you. That is exactly why the finding matters, but it is worth stating that a combination which underperforms an oracle-selected best performer may still beat the realistic alternative of picking one and hoping. The honest reading is not that human-AI teams are useless. It is that the pairing has to be designed against a known division of competence, and that undesigned pairing is worse than either party alone.
Most of the underlying studies are laboratory or bounded-task experiments, with short horizons and clear right answers. Real organisational decisions are longer, more ambiguous, more political, and rarely scored. And none of this measures the second-order effect: what happens to a person's decision-making capability after two years of supervising a machine rather than deciding. That question is taken up in how humans learn with AI.
The SuperSkills interpretation
Human in the loop is not a design. It is a reassurance, and it has become the phrase organisations reach for at precisely the moment they would otherwise have to think. Putting a person at the end of an automated process, with no time budget, no authority to stop it, no stated basis on which they would disagree, and no consequence if they simply approve, does not produce oversight. It produces a signature. The meta-analysis is the empirical case for something I have argued for years from practice: the presence of a human does not improve a decision, and a human who has been positioned to rubber-stamp will make the system worse than either party alone, because they add latency and the appearance of scrutiny without the substance.
Which is why the useful question is not whether a human is involved but where. Reviewing at the end is the weakest position available: the framing is already set, the options have been narrowed, anchoring has happened, and the effort required to reopen the question is far higher than the effort required to approve. Being present at the start is a different job entirely, because that is where the problem is defined, the constraints are set, the success criteria are chosen and the alternatives that will never be generated are quietly excluded. That is the principle I call Human at the Start, and the evidence on framing effects and automation bias says it is not a stylistic preference. It is the only point in the process where a human still has leverage that is cheaper to exercise than to skip.
Read the Vaccaro result carefully and it yields three conditions under which pairing genuinely helps, which is more useful than the headline. Pairing helps when the human is genuinely better at the task, so the machine augments rather than leads. It helps when the task is generative rather than evaluative, because content creation showed the gains while decision-making showed the losses. And it helps when the human's contribution is defined as something other than approval, such as setting the problem, supplying context the model cannot have, or making the value trade-off the model has no standing to make. Where none of those three holds, you are not designing a partnership. You are adding a person to a process for comfort, and paying for it in accuracy.
The commercial consequence follows directly. If verification is the human's job, then verification is the skilled work and should be resourced, trained and paid as such. Almost no organisation does this, because checking looks like an administrative act and producing looks like a professional one. I call that mispricing the verifier's discount, and the jagged-frontier result is what it costs.
Deciding who decides: a working allocation
Before a class of decision reaches a workflow, answer four questions in writing. The point of writing them down is not bureaucracy; it is that an organisation which has never written them has already answered them by default.
- Who is better at this, honestly? Not who should be. On this specific task, with this data, is the model more accurate than your people, less accurate, or unknown? Unknown is a legitimate answer and it changes the design, because it means you are running an experiment and should measure it as one.
- What is the human actually adding? Name it: framing, context, the value judgement, accountability, or catching a specific known failure mode. If the honest answer is "confidence", remove them from the loop and put them at the start instead.
- What would make a person disagree, and can they? Specify in advance the conditions under which the output should be rejected, and confirm the person has the time, the standing and the authority to reject it. Oversight without stop-work authority is theatre.
- How would we know this is going wrong? Approval rates near a hundred percent are a signal, not a success. Sample decisions and score them independently, because a process where nobody ever disagrees is indistinguishable from a process where nobody is looking.
For decisions taken by agents rather than assistants, where the system plans and acts rather than advises, the allocation question becomes sharper still and is developed separately in AI agents and human judgement.
What this looks like in practice
A credit team runs a model that is measurably better than its analysts at scoring standard applications, and worse at the unusual ones. The wrong design is a human reviewing every case, which adds delay and drags the good decisions towards the human's lower accuracy. The right design routes the standard cases to the model with sampled audit, and routes the unusual ones to a person before the model has framed them. Same technology, opposite performance, and the difference is one page of thinking nobody had time for.
A clinical team uses an AI triage tool. The failure to design against is not the model being wrong; it is the model being right often enough that questioning it starts to feel like obstruction. That is the automation-complacency finding, and the countermeasure is structural rather than attitudinal: a stated disagreement rate that is expected to be non-zero, protected time to examine flagged cases, and no professional cost for overriding.
A leadership team receives an AI-generated market analysis, discusses it for forty minutes, and approves the recommendation. Nobody asks what question was put to the model, what it was not given, or which options it never generated. The meeting felt rigorous. Every input to the decision was set before anyone in the room was involved.
What to do
Stop saying human in the loop and start saying who, where and with what authority. The phrase has become a way of not deciding. Replace it in your policies with a named person, a named point in the process, and a stated basis for disagreement.
Move the human upstream. Framing, constraints and success criteria are where a person still has leverage. Review at the end is where they have the least, and it is where almost every organisation has put them.
Establish which of you is better, task by task, and write it down. This is the single highest-return piece of work available and it is nearly always skipped, because it requires admitting that on some tasks the machine is better and on others your people are, and both admissions are politically awkward.
Resource verification like the skilled work it is. Give it time, training, seniority and status, or accept that you have bought a process which fails exactly when it matters, on the tasks that sit outside the model's competence.
Measure disagreement. Track how often humans override, and investigate when the rate approaches zero. A hundred percent approval is not evidence of a good model. It is evidence of an unmeasured one.
Development of the idea
I argued in Entrepreneur UK in July 2026 that AI does not create bad decisions, it exposes them faster, which is the same point from the organisational side: AI reveals whether a decision process was ever a process. The accountability form of the argument, including the distinction between Human at the Start, human in the loop and human at the end, is set out in the European Business Review piece on accountability gaps in leadership decisions (21 August 2026). In the Observer in July 2026 I argued that judgement, not model-building, would define the next power class. The framework is developed in SuperSkills (Kogan Page, 2026).
Key research and primary sources
- Vaccaro, M., Almaatouq, A. and Malone, T. (2024). When combinations of humans and AI are useful: a systematic review and meta-analysis. Nature Human Behaviour, 8, 2293-2303.
- Dell'Acqua, F. et al. (2023). Navigating the Jagged Technological Frontier. Harvard Business School and BCG working paper.
- Parasuraman, R. and Manzey, D. H. (2010). Complacency and Bias in Human Use of Automation. Human Factors, 52(3).
- Dietvorst, B. J., Simmons, J. P. and Massey, C. (2015). Algorithm Aversion: People Erroneously Avoid Algorithms After Seeing Them Err. Journal of Experimental Psychology: General, 144(1).
- Logg, J. M., Minson, J. A. and Moore, D. A. (2019). Algorithm Appreciation: People Prefer Algorithmic to Human Judgment. Organizational Behavior and Human Decision Processes, 151, 90-103.
- Brynjolfsson, E., Li, D. and Raymond, L. (2023). Generative AI at Work. NBER Working Paper 31161; published in the Quarterly Journal of Economics, 2025.
- Lee, H.-P. et al. (2025). The Impact of Generative AI on Critical Thinking. Microsoft Research and Carnegie Mellon, CHI 2025.
Related SuperSkills research
On where the human belongs in the process, Human at the Start and AI agents and human judgement. On the underlying capability, AI and human judgement and decision quality in the AI era. On the mispricing of verification, the verifier's discount. On the organisational conditions, drift versus design and AI workforce strategy. The leadership framing is how leaders should respond to AI.
About this research
Rahim Hirji is the author of SuperSkills: The Seven Human Skills for the Age of AI (Kogan Page, 2026) and the founder of The SuperSkills Intelligence Company. This work draws on research across more than 200 organisations in 30 countries over seven years. Findings are attributed to the studies that produced them and kept separate from the interpretation, which is the author's. Automation bias, automation complacency, algorithm aversion and algorithm appreciation are established concepts from the research literature and are not his. Human at the Start, drift versus design and the verifier's discount are part of the SuperSkills lexicon. This is a living reference, reviewed and updated as significant new evidence appears.
Cite this
Hirji, R. (2026). Human and AI decision making. The SuperSkills Intelligence Company. Last reviewed 26 August 2026. thesuperskills.com/research/human-ai-decision-making