← Research
Research

AI agents and human judgement

When the machine stops answering and starts acting, the human's job is no longer to do the work. It is to decide what to delegate, and to own the result.

Last reviewed: 25 August 2026

AI agents plan, use tools and act across many steps toward a goal, with little or no supervision. That changes the human role from doing to delegating and owning. This page sets out that shift, the delegation boundary map, and who is accountable when an agent is wrong. Developed in Rahim Hirji's work, including the Box of Amazing essay The Agents Are Here.

Questions this page answersAll 780 questions this research covers

The arrival of AI agents changes who does the work. Calling it another productivity upgrade misses that entirely. A chatbot answers a question and hands the result back to a person; an agent takes a goal, plans the steps, uses tools and acts, often across systems, often without anyone watching each move. When the machine acts rather than answers, the human's job moves up a level: from doing the task to deciding what may be delegated, setting the boundaries, and owning the outcome. The danger in the agentic shift has little to do with unsafe agents. Accountability and capability both drift away , the work still gets done, but no one made the decision and no one is building the judgement to make the next one. The organisations that handle agents well will be the ones that decide, deliberately, where the human belongs.

What an agent is, and why it changes the question#

The useful distinction is simple. A chatbot responds. An agent acts: it decomposes a goal into steps, calls tools and other systems, adapts as it goes, and completes a sequence with limited human involvement. The moment this became visible to everyone, as Rahim Hirji wrote in The Agents Are Here, was less about any single product and more about a shift in posture: software that does things on your behalf while you watch from the sidelines. The right response, he argued, is not the usual pair of lenses, excitement or fear, but a third: what does this do to human capability, and what should we keep firmly human? That is the question this page answers.

The human's new job#

With agents, the human moves from operator to something closer to a principal with four roles. Intent-setter: defining what a good outcome is before the agent runs. Delegator: deciding what the agent may and may not do. Verifier: checking the output, and separately whether the agent stayed inside its boundaries. Accountable owner: the named person who can explain the result and had the standing to stop it. This is Human at the Start applied to agents. The consequential judgement is made at the framing, before the agent acts, and owned at the end. Everything the agent does in between is execution, however autonomous it looks.

The delegation boundary map#

The practical instrument is a map of what the human keeps and what the agent may take, decision by decision rather than task by task. For any consequential piece of work:

The map is not a rule that a human must touch everything, which would defeat the point of agents. It is a decision about where the human touch has to be, made in advance, so that autonomy is granted deliberately rather than by default.

Human in the loop, or human on the hook#

Boards reach for "we keep a human in the loop", and with agents the phrase breaks. When an agent acts quickly and at scale, a person placed in the middle to review each step is either overwhelmed or reduced to a rubber stamp, and that is where oversight is weakest. As Rahim Hirji sets out in the accountability work, drawn from a landmark AI-denial case and European law, meaningful oversight requires someone with the authority and competence to change the decision, and producing an automated recommendation can itself be the decision. With agents, the loop is not the answer. A named human at the start, and a named human on the hook at the end, is. If no name can be attached to an agent's work, the work has not been delegated; it has been abandoned.

Automation complacency, multiplied#

The risk the loop is meant to catch gets worse with agents, not better. Parasuraman and Manzey, reviewing decades of research across aviation, medicine and the military, found that people under-question confident automated output, that this affects experts as much as novices, that it cannot be trained away, and that it worsens under load and when attention is split across tasks. An agent running many actions at once is a split-attention, high-load situation of the worst kind. The more the agent does, the less any human scrutinises, and the Harvard and BCG jagged-frontier experiment showed the cost: people who trusted AI beyond its competence performed worse than those with no AI at all. Complacency is not a character flaw here; it is the predictable result of designing humans into the weakest position.

What agents do to early-career development#

There is a slower, more serious effect. If junior staff move straight to supervising agents rather than doing the underlying work, they never accumulate the repetitions that build judgement. The output looks senior; the capability is not. This is synthetic seniority and the missing rungs, accelerated, and across an organisation it compounds into capability debt: an operation that runs smoothly on agents until a decision arrives that the agent cannot make and no human in the room has been trained to. Agents make this cheaper to ignore and more expensive when the bill arrives.

What leaders should do#

Decide the delegation boundaries before deploying an agent, not after an incident: what it may do, what it may never do, and what would trigger a human override. Name an accountable owner for every consequential agent workflow, in writing, who can explain the outcome without reference to the tool. Treat verification of agents as real, skilled work, checking boundaries and reasoning, not just outputs, because in an agentic operation the checking is the judgement. Protect the reps: keep deliberate practice in the system so people still build the capability the agents are now exercising. And watch the development curve alongside the throughput curve, because an organisation can look more productive and grow less capable at the same time. The point of agents is to raise what people can get done. The job of leadership is to make sure it does not lower what they can decide.

Key sources

Key research and primary sources

Every source below has been fetched and confirmed. The graded versions, including what each does not support, are in the evidence base.

This extends the account of judgement to agents: Human at the Start, AI and human judgement, capability debt, synthetic seniority and drift versus design. On what cannot be delegated to an agent at all, see what stays human. The evidence on pairing humans with AI for decisions is reviewed in human and AI decision making. The operational version, stage by stage with a downloadable grid, is the Delegation Boundary Map. On what the law now requires of oversight, meaningful human oversight. The position that follows from this, put simply, is that human in the loop is not a safeguard. See should I let an agent act on my behalf.

About this research#

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. This work draws on research across more than 200 organisations in 30 countries over seven years, and on his ongoing writing on the agentic shift. Findings are attributed to their sources and kept separate from the interpretation and frameworks, which are the author's. This is a living reference, and a fast-moving one: it is reviewed at least every 90 days as agent capability changes.

How this research works  ·  Reviewed quarterly  ·  Found an error? Tell me and it is corrected on the page.

Cite this

Hirji, R. (2026). AI agents and human judgement. The SuperSkills Intelligence Company. Last reviewed 25 August 2026. thesuperskills.com/research/ai-agents-and-human-judgement

Questions answered on this page

What is an AI agent, and how is it different from a chatbot?

A chatbot responds; an agent acts. An AI agent takes a goal, plans the steps, uses tools, and carries out a sequence of actions across systems with limited or no supervision, rather than returning a single answer for a human to use. That shift, from producing output to taking action, is what moves the human's job from doing the work to deciding what to delegate and owning the result.

Who is accountable when an AI agent makes a mistake?

A named human, or the decision was not ready to be delegated. An agent acting unsupervised makes accountability easy to lose: the work happens, but no person can honestly say they made the call. Rahim Hirji's answer is to name an owner before the agent runs, someone who defined what the agent may do and what would stop it, and who can explain and answer for the outcome. If no name can be attached, the task is not ready for an agent. See AI accountability and Human at the Start.

Should you keep a human in the loop with AI agents?

Human in the loop is the weakest of the options and, with fast agents acting at scale, often impossible in practice. The stronger design is a human at the start, who sets the intent, the boundaries and the override conditions before the agent runs, and a human at the end who owns the outcome. The loop in the middle will not save you when the agent has already acted a thousand times.

How does agentic AI affect junior employees?

It risks accelerating capability debt. If juniors move straight to managing agents rather than doing the underlying work, they never build the judgement the senior role will demand, and the output looks senior while the capability is not. This is synthetic seniority and the missing rungs, sped up. The fix is to keep deliberate practice in the system even as agents take the tasks.

In this hub

Judgement, oversight and accountability

Who decides, who checks, and who is answerable when the machine was involved.

Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Before you deploy agents is the cheapest moment to decide this, and the only one where you are not deciding under pressure. Settling what an agent may do unasked, and who answers for it, is the work. AI advisory for CEOs and boards.

Agents are where this stops being a thought experiment. There is the AI agents and accountability version, and the full range of topics and audiences.

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.