- Should government AI agents act on citizens' behalf, and under whose authority?
- How do you supervise something that does not wait for you?
- What permissions should an AI coding agent have on live files?
- Can an AI agent give out my address or agree a price for me?
- Should an AI agent be allowed to spend money without approval?
- Can an AI agent have delegated authority?
Sometimes, and trust has little to do with it. The question is whether you have specified the boundary, because an agent without a stated boundary is abdication with a progress bar.
The answer, in one line
Sometimes, and the useful question is not whether you trust it. It is what the worst thing is that it can do before a human sees it, and whether you can live with that.
The distinction that matters is simple and almost nobody makes it. A system that answers gives you something to accept or reject. A system that acts has already done it. Every oversight model in common use assumes a pause for review, and agents remove the pause.
The question that replaces "do I trust it?"#
What is the worst thing this can do before a human sees it, and can I live with that? Answer that and the trust question resolves itself. Leave it unanswered and no amount of confidence in the model helps you, because you have not bounded the downside.
Four things to fix before granting autonomy#
1 · Reversibility. Sort the actions the agent can take into reversible, expensive to reverse, and irreversible. Grant autonomy freely in the first category, reluctantly in the second, and not at all in the third without a stop. This does more work than any accuracy estimate, because it bounds the loss rather than the probability.
2 · Blast radius. Not what the agent does, but how far the consequence travels. Sending one email is small. Sending one email to a client list is not, and the action is identical.
3 · The stop. Can you halt it mid-sequence, and does halting leave things in a safe state or a broken one? Article 14 of the EU AI Act requires exactly this for high-risk systems: intervention or interruption bringing the system to a halt in a safe state. Most consumer and internal agent deployments have not been designed for it.
4 · The reconstruction test. If this goes wrong, can you reconstruct why the agent did what it did? If the reasoning exists only inside a sequence of model calls, it has already gone, and you will be explaining an outcome you cannot account for.
Why the evidence points at the front rather than the back#
The Vaccaro meta-analysis of 106 experiments found human-AI combinations underperforming the better party alone, with losses concentrated in decision tasks, where a human judges whether a system was right, and gains in creation tasks, where the pair produces something together.
Agents make the decision task worse in the one way that matters: they perform it at machine speed, in volume, without the human present. If review after the fact was already the weakest available position, reviewing after the fact and after execution is weaker still.
Which means the human contribution has to move to where it counts: the problem, the intent, the constraints and the rejection criteria, all before anything runs. That is Human at the Start, and with agents it stops being a preference and becomes the only place the human can meaningfully be.
What is genuinely uncertain#
Most of it. There is no equivalent of the Vaccaro analysis for agentic systems, no field evidence on agent oversight failures at scale, and no established practice for what adequate autonomy limits look like. Agent capability is also moving faster than any other part of this field, so anything written now dates quickly.
Anyone offering confident guidance on agent delegation, including this page, is reasoning from adjacent evidence. The difference is whether they say so.
Agents make the argument unavoidable#
Agents are the point at which the argument this research has been making becomes unavoidable rather than advisory. When a system answers, you can compensate for a weak oversight design by being careful at the end. When a system acts, there is no end to be careful at. The decision was made when you set the boundary, or it was not made at all.
There is a second consequence, quieter and worse. Agents absorb exactly the sequences of small tasks that used to constitute learning a job: chasing the thing, checking the thing, noticing the anomaly, following it up. Those look like overhead and they are where judgement is formed. An organisation that hands them wholesale to agents has not just automated coordination, it has removed the last visible route by which anyone learned how the work actually holds together. See the missed reps.
A working rule#
- Let agents act freely on reversible, low-radius tasks. Most of the value is here and most of the anxiety is not.
- Require a human at the boundary of any irreversible or externally visible action. Not a review of the output, a decision before it runs.
- Write the limits down. Spending caps, recipient lists, systems it may touch, actions it may never take. Unwritten limits are not limits.
- Log the reasoning, not just the actions. Otherwise the reconstruction test fails at the moment you need it.
- Keep doing enough of the work yourself to notice when the agent is wrong. The capability that lets you set a good boundary is the same capability the agent is removing.
29 and 30 September 2026: agents that never switch off, and a standing permission#
Two products launched on 29 September make this page’s question a consumer one. OpenAI’s dots are “always-on agents” with “their own cloud computer” that “can work towards your goals 24/7”. The company describes tiers of control. Background research uses tools “restricted to be read-only”. Actions that could affect accounts or share information pass an automatic review that decides “what work can proceed, what needs approval, and what you must do yourself”. A third tier is reserved: “Certain sensitive tasks, such as changing a password, always stay with you.” One further sentence hands the judgement back: “Dots can still make mistakes, so always review consequential work.” Meta’s Muse for Small Business promises that “nothing publishes, sends, or spends without your approval.” The first reported test of an approval design came a day later. Caleb Kinchlow at TechRepublic described a Facebook Marketplace seller’s account of Muse negotiating a sale, agreeing a lower price, giving a buyer the pickup location and telling him on arrival that the seller was there. The seller had chosen “Allow Always” when setting the agent up, believing it covered messages and that larger actions would still be put to him. Meta said there had been “no breach of privacy controls”. It is one account, and both parties may be right: the control worked as built, and the person did not understand what he had granted. Blast radius, the second check on this page, is what the case turns on. The action was a message and the consequence was a stranger at the door. A standing permission is a delegation of authority and needs what any delegation needs: a scope in words the grantor would recognise, a limit set by consequence and not by type of action, and an expiry. Shane Savitsky at Axios collected polls suggesting most people sense this: in a Thales survey “only 13% of respondents would let an ‘AI helper’ read their emails” and 7 per cent would let one move money, though the article gives no sample sizes. Banks say the same of themselves. Brian Moynihan of Bank of America, quoted by Nathan Place at American Banker on 29 September, said the risk “largely is around agents just left to operate”, and added: “We just don’t do that.”
Related SuperSkills research#
On judgement with agents, AI agents and human judgement. On the stage-by-stage version, the Delegation Boundary Map. On why review fails, human in the loop is not a safeguard. On the legal duty, meaningful human oversight. On the override rule, when should I override AI.
Key sources
- Vaccaro, M., Almaatouq, A. and Malone, T. (2024). When combinations of humans and AI are useful. Nature Human Behaviour, 8.
- Bainbridge, L. (1983). Ironies of Automation. Automatica, 19(6).
- Parasuraman, R. and Riley, V. (1997). Humans and Automation: Use, Misuse, Disuse, Abuse.
- Article 14, Human Oversight, Regulation (EU) 2024/1689.
About this research#
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. Agent-specific evidence is thin and this page reasons from adjacent findings, which it states above. Not legal advice. On a 90-day review cycle, because this is the fastest-moving question on the site.
Essay · SS-2026-057
Hirji, R. (2026). Should I let an AI agent act on my behalf?. The SuperSkills evidence base, SS-2026-057. https://thesuperskills.com/research/should-i-let-an-ai-agent-act-on-my-behalf. Last reviewed 2 October 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work