← Research
Research

Human and AI decision making

On average, putting a human and an AI together produces a worse decision than the better of the two alone. That finding should reorganise how you design oversight.

Last reviewed: 26 August 2026

How should humans and AI make decisions together? The evidence does not support the answer everyone gives. This page sets out what the meta-analysis actually found, and how Rahim Hirji thinks decisions should be allocated between people and machines.

Question this page answersAll 616 questions this research covers

Here is the finding that should reorganise how organisations think about oversight, and which almost nobody in the field has absorbed. Across 106 experimental studies and 370 effect sizes, putting a human and an AI together produced decisions that were on average worse than the better of the two working alone. Not worse than the human. Worse than whichever of the two was better at that task. The losses concentrated specifically in decision-making, while content creation showed real gains. Keep a human in the loop is the answer everyone gives to this, and the evidence does not support it. What survives is narrower: pairing helps when the human is better than the machine at the task, and hurts when the machine is better and the human overrides it or fails to catch it. Which means the whole discipline lives upstream, in working out which of you is actually better at what, before the decision arrives.

What the meta-analysis found#

The anchor result is a preregistered systematic review and meta-analysis by Vaccaro, Almaatouq and Malone, published in Nature Human Behaviour in 2024. They gathered 106 experimental studies reporting 370 effect sizes and asked a simple question: does a human working with an AI outperform the best of a human alone or an AI alone? On average, no. The combination performed significantly worse, with an effect size of Hedges' g of -0.23. The headline is arresting, but the structure underneath it is what matters. Losses were concentrated in decision-making tasks and gains were significantly larger in content-creation tasks. And the direction was predictable: where humans alone outperformed the AI, the combination gained; where the AI alone outperformed humans, the combination lost. Pairing did not average the two. It frequently dragged the better performer down towards the worse one.

That mechanism has a long-established name. Parasuraman and Manzey, reviewing decades of work across aviation, medicine and the military in 2010, described automation bias and complacency: the tendency to under-question automated advice, present in novices and experts alike, resistant to training, and worse under time pressure and high workload. Ignorance has nothing to do with it. This is what attention does when a competent system is doing the monitoring for you.

The most quoted demonstration in a professional setting is the jagged technological frontier. In 2023, Dell'Acqua and colleagues, working with Boston Consulting Group and researchers at Harvard, MIT and Wharton, gave 758 consultants access to GPT-4. Inside the model's competence, AI-assisted consultants were dramatically better and faster. On a task deliberately designed to sit just outside it, consultants using AI performed worse than consultants with no AI at all. The frontier is jagged rather than smooth, which is the difficult part: you cannot infer from a model's brilliance on one task that it is competent on an adjacent one, and the confidence of its output tells you nothing about which side of the line you are on.

Two findings from the trust literature explain why calibration is so hard, and they appear to contradict each other until you look at when each applies. Dietvorst, Simmons and Massey, writing in the Journal of Experimental Psychology: General in 2015, described algorithm aversion: after seeing an algorithm err, people abandon it, even when it demonstrably outperforms them and even when they have just watched their own worse performance. Logg, Minson and Moore, in 2019, described the opposite tendency, algorithm appreciation: for many estimates and forecasts, people weight algorithmic advice more heavily than human advice, with experts in the domain being the notable exception. Put together, they show trust moving for reasons unrelated to accuracy. Too much trust or too little turns out to be the wrong axis. And a policy instructing people to use judgement cannot fix a miscalibration it never diagnoses.

Finally, the gains are real and they are unevenly distributed. Brynjolfsson, Li and Raymond, studying 5,172 customer-support agents, found AI assistance raised productivity by fifteen percent on average, thirty percent for the newest and least experienced staff, and almost nothing for the most skilled. And the 2025 Microsoft Research and Carnegie Mellon survey of 319 knowledge workers found that the higher a worker's confidence in the AI, the less critical thinking they reported applying, with effort shifting from producing the work to verifying the output.

The studies stop at June 2023#

The meta-analysis covers studies published between January 2020 and June 2023, which means most of the underlying work predates the current generation of frontier models. If model capability has risen since, the balance of "who is better at this task" has shifted in the machine's favour, and the practical implication changes: more tasks fall into the category where human intervention subtracts rather than adds. That does not weaken the finding. It sharpens the warning, because it means more of the decisions where a human is nominally in the loop are decisions where the human is nominally in the way.

There is also a fairness point about the benchmark. "Worse than the best of either alone" is a demanding comparison, because in real settings you rarely know in advance which of the two is better at the specific task in front of you. That uncertainty is why the finding matters, though a combination which underperforms an oracle-selected best performer may still beat the realistic alternative of picking one and hoping. So the reading to take away is narrower than "human-AI teams are useless". Pairing has to be designed against a known division of competence. Undesigned, it comes out worse than either party alone.

Most of the underlying studies are laboratory or bounded-task experiments, with short horizons and clear right answers. Real organisational decisions are longer, more ambiguous, more political, and rarely scored. And none of this measures the second-order effect: what happens to a person's decision-making capability after two years of supervising a machine rather than deciding. That question is taken up in how humans learn with AI.

Human in the loop describes nothing#

Human in the loop is a reassurance rather than a design. It has become the phrase organisations reach for at the moment they would otherwise have to think. Putting a person at the end of an automated process, with no time budget, no authority to stop it, no stated basis on which they would disagree, and no consequence if they simply approve, does not produce oversight. It produces a signature. The meta-analysis is the empirical case for something I have argued for years from practice: the presence of a human does not improve a decision, and a human who has been positioned to rubber-stamp will make the system worse than either party alone, because they add latency and the appearance of scrutiny without the substance.

So the useful question is where the human sits. Reviewing at the end is the weakest position available: the framing is already set, the options have been narrowed, anchoring has happened, and the effort required to reopen the question is far higher than the effort required to approve. Being present at the start is a different job entirely, because that is where the problem is defined, the constraints are set, the success criteria are chosen and the alternatives that will never be generated are excluded. That is the principle I call Human at the Start, and the evidence on framing effects and automation bias makes it more than a stylistic preference. The start is the only point in the process where a human has leverage that is cheaper to exercise than to skip.

Read the Vaccaro result carefully and it yields three conditions under which pairing genuinely helps, which is more useful than the headline. Pairing helps when the human is better at the task, so the machine augments rather than leads. It helps when the task is generative rather than evaluative, because content creation showed the gains while decision-making showed the losses. And it helps when the human's contribution is defined as something other than approval, such as setting the problem, supplying context the model cannot have, or making the value trade-off the model has no standing to make. Where none of those three holds, you are not designing a partnership. You are adding a person to a process for comfort, and paying for it in accuracy.

The commercial consequence follows directly. If verification is the human's job, then verification is the skilled work and should be resourced, trained and paid as such. Almost no organisation does this, because checking looks like an administrative act and producing looks like a professional one. I call that mispricing the verifier's discount, and the jagged-frontier result is what it costs.

Deciding who decides: a working allocation#

Before a class of decision reaches a workflow, answer four questions in writing. The point of writing them down is not bureaucracy; it is that an organisation which has never written them has already answered them by default.

For decisions taken by agents rather than assistants, where the system plans and acts rather than advises, the allocation question becomes sharper still and is developed separately in AI agents and human judgement.

What this looks like in practice#

A credit team runs a model that is measurably better than its analysts at scoring standard applications, and worse at the unusual ones. The wrong design is a human reviewing every case, which adds delay and drags the good decisions towards the human's lower accuracy. The right design routes the standard cases to the model with sampled audit, and routes the unusual ones to a person before the model has framed them. Same technology, opposite performance, and the difference is one page of thinking nobody had time for.

A clinical team uses an AI triage tool. The failure to design against is not the model being wrong; it is the model being right often enough that questioning it starts to feel like obstruction. That is the automation-complacency finding, and the countermeasure is structural rather than attitudinal: a stated disagreement rate that is expected to be non-zero, protected time to examine flagged cases, and no professional cost for overriding.

A leadership team receives an AI-generated market analysis, discusses it for forty minutes, and approves the recommendation. Nobody asks what question was put to the model, what it was not given, or which options it never generated. The meeting felt rigorous. Every input to the decision was set before anyone in the room was involved.

Say who, where, and with what authority#

Stop saying human in the loop and start saying who, where and with what authority. The phrase has become a way of not deciding. Replace it in your policies with a named person, a named point in the process, and a stated basis for disagreement.

Move the human upstream. Framing, constraints and success criteria are where a person still has leverage. Review at the end is where they have the least. It is where almost every organisation has put them.

Establish which of you is better, task by task, and write it down. This is the single highest-return piece of work available and it is nearly always skipped, because it requires admitting that on some tasks the machine is better and on others your people are, and both admissions are politically awkward.

Resource verification like the skilled work it is. Give it time, training, seniority and status, or accept that you have bought a process which fails exactly when it matters, on the tasks that sit outside the model's competence.

Measure disagreement. Track how often humans override, and investigate when the rate approaches zero. A hundred percent approval evidences an unmeasured process rather than a good model.

Development of the idea#

I argued in Entrepreneur UK in July 2026 that AI does not create bad decisions, it exposes them faster, which is the same point from the organisational side: AI reveals whether a decision process was ever a process. The accountability form of the argument, including the distinction between Human at the Start, human in the loop and human at the end, is set out in the European Business Review piece on accountability gaps in leadership decisions (21 August 2026). In the Observer in July 2026 I argued that judgement, not model-building, would define the next power class. The argument is set out at length in SuperSkills (Kogan Page, 2026).

What regulators have already decided about oversight#

The argument on this page has a legal counterpart, and it has been settled more clearly in European regulation than in most corporate policy. The test the courts apply is competence and an actual decision, not presence.

In 2023 the Amsterdam Court of Appeal ruled against Uber in a case brought by drivers deactivated after fraud flags. What the court examined was not the presence of a review step but whether the reviewer's qualifications and knowledge could be evidenced. Uber could not say who had decided or what they knew. In 2024 Italy's data protection authority, the Garante, fined Foodinho, the Glovo subsidiary, five million euros, a second offence after a 2.6 million euro fine in 2021. The remedy is the interesting part: reviewers must be adeguatamente formati, adequately trained. Competence, not attendance, became the standard.

In August 2026 the Dutch data protection authority fined Uber roughly 825 million euros over driver deactivations decided without human assessment, a decision Uber is appealing. The finding rests on an absence rather than an act: scoring drivers algorithmically remained lawful, and deciding their livelihood without a person did not.

There is a lesson here for anyone writing an AI policy. "A human reviews the output" is not a control that survives contact with a regulator, and increasingly not one that survives contact with a court. Who, with what training, deciding what, on what basis, with what authority to refuse: those are the questions being asked, and an organisation that cannot answer them has documented oversight rather than exercised it.

Nobody can tell you who benefits#

There is a further complication for anyone designing an oversight policy. It is the most awkward finding in this whole literature. Yu and colleagues, publishing in Nature Medicine in 2024, gave 140 radiologists AI assistance across fifteen chest X-ray tasks, roughly 5,190 observations, randomised, with statistical treatment designed to separate genuine individual differences from noise.

The effects diverged sharply. AI assistance helped some radiologists substantially and made others measurably worse. And nothing predicted which: not years of experience, not subspecialty, not prior familiarity with AI. Lower performers did not reliably gain, which is the assumption most deployment plans rest on.

The implication is uncomfortable, so let me put it plainly. Rolling a tool out to everyone in a professional group will help some of them and harm others, and at present there is no way to know in advance who is in which group. That is an argument for measuring individual effects rather than assuming an average. It is a further reason why "we have kept a human in the loop" describes a hope rather than a control.

What the automation literature settled decades ago#

Almost every finding above has a precedent in human-factors research, and reading it is the cheapest available upgrade to how an organisation designs oversight. Parasuraman and Riley, writing in Human Factors in 1997, set out a taxonomy that is still sharper than the vocabulary in current AI policy: use, misuse, disuse and abuse. Misuse is over-reliance. Disuse is unwarranted rejection, which is a real and separate failure. Abuse is deployment by designers in ways that ignore the human consequences. Collapsing all three into "over-reliance", as most AI guidance does, loses the distinctions that determine what you should actually do.

Skitka, Mosier and Burdick demonstrated automation bias experimentally in 1999 and split it into two kinds of error, and the split matters more than the phenomenon. Commission errors are acting on a wrong recommendation. Omission errors are missing something the system did not flag. Nearly every oversight process is designed to catch the first. The second is invisible by construction, because nothing appears on the screen to check. That is the failure mode that accumulates.

The most uncomfortable finding is Dzindolet and colleagues, in 2003. Trust mediates reliance, as you would expect. But they also found that explaining why an automated aid might err increased reliance on it, even when that restored trust was not warranted. That is a direct warning about explainability as a safety measure: telling people how a model can fail may make them trust it more rather than less. If your governance rests on transparency producing appropriate scepticism, this study says the mechanism can run backwards.

For deciding how far to automate rather than whether, Parasuraman, Sheridan and Wickens set out a four-stage, ten-level model in 2000 that remains more rigorous than most current thinking about what to delegate to an agent. And on the practical side, Daugherty and Wilson's Human + Machine names the hybrid roles in what they call the missing middle, where humans train, explain and sustain machine systems. The NIST AI Risk Management Framework gives this a governance vocabulary a board will recognise, though its weakness is instructive: it can be satisfied procedurally by an organisation that documents oversight without exercising it.

Key research and primary sources

On where the human belongs in the process, Human at the Start and AI agents and human judgement. On the underlying capability, AI and human judgement and decision quality in the AI era. On the mispricing of verification, the verifier's discount. On the organisational conditions, drift versus design and AI workforce strategy. The leadership framing is how leaders should respond to AI. The graded evidence is in the evidence base. On the underlying tendency, automation bias; on the meta-analytic baseline, what is human-AI collaboration? The operational version, stage by stage with a downloadable grid, is the Delegation Boundary Map. On who carries the verification duty, who owns verification when AI does the work. The position that follows from this, put simply, is that human in the loop is not a safeguard.

About this research#

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. This work draws on research across more than 200 organisations in 30 countries over seven years. Findings are attributed to the studies that produced them and kept separate from the interpretation, which is the author's. Automation bias, automation complacency, algorithm aversion and algorithm appreciation are established concepts from the research literature and are not his. Human at the Start, drift versus design and the verifier's discount are part of the SuperSkills lexicon. This is a living reference, reviewed and updated as significant new evidence appears.

How this research works  ·  Reviewed quarterly  ·  Found an error? Tell me and it is corrected on the page.

Cite this

Hirji, R. (2026). Human and AI decision making. The SuperSkills Intelligence Company. Last reviewed 26 August 2026. thesuperskills.com/research/human-ai-decision-making

Questions answered on this page

How should humans and AI make decisions together?

By deciding in advance which of the two is better at the specific task, rather than routinely adding a person at the end. A 2024 meta-analysis in Nature Human Behaviour by Vaccaro, Almaatouq and Malone, covering 106 studies and 370 effect sizes, found human-AI combinations performed significantly worse than the better of human or AI alone, with losses concentrated in decision-making and gains in content creation. Pairing helped where humans outperformed the AI and hurt where the AI outperformed humans. Rahim Hirji argues the human's leverage is at the start, in framing and constraints, not at the end in review.

Does keeping a human in the loop improve AI decisions?

Not by itself, and often the opposite. A person placed at the end of an automated process, with no time budget, no stated basis for disagreement and no authority to stop it, produces a signature rather than oversight, and adds latency and the appearance of scrutiny without the substance. The meta-analytic evidence shows undesigned pairing performing worse than either party alone. What matters is where the human sits, what they are adding, and whether they can actually say no.

When does human-AI collaboration actually help?

Three conditions, read from the Vaccaro meta-analysis. First, when the human is genuinely better at the task, so the machine augments rather than leads. Second, when the task is generative rather than evaluative, since content creation showed gains while decision-making showed losses. Third, when the human's contribution is defined as something other than approval, such as framing the problem, supplying context the model cannot have, or making a value trade-off the model has no standing to make.

How do you tell if AI oversight is real or theatre?

Measure disagreement. If the human approval rate is close to a hundred percent, you have an unmeasured process rather than a good model. Ask four questions in writing before a class of decision reaches a workflow: who is genuinely better at this task, what is the human actually adding, what would make a person disagree and can they, and how would we know this is going wrong. Oversight without stop-work authority is theatre.

In this hub

Judgement, oversight and accountability

Who decides, who checks, and who is answerable when the machine was involved.

The work

Where the writing comes from.

These essays draw on research across more than 200 organisations in 30 countries. See the wider body of work, or bring it into your organisation.

All research →
Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory and coaching  ·  Enquire