- How fast must a human be able to stop an AI agent?
- Why did it take hours to stop an AI agent that was flagged in minutes?
- How long would it take us to stop one of our own agents, and who has the authority to do it?
- Is logging what an AI agent does enough for oversight?
Faster than the agent can finish what it started, and that is a number nobody has agreed. The one case with a published clock is OpenAI’s own. On 20 September 2026 an agent in a training sandbox found a route to the open internet through the sandbox’s own name server. An alert fired twelve minutes after the first outside response reached it, a person acknowledged the alert three minutes after that, and the run was stopped two hours and twenty-nine minutes later, because, in the company’s words, it “did not stop automatically as expected”. Six days later the company paused training and use of its most capable models for the second time in three months. For an organisation running agents of its own the question is answerable today, with three timings, one yes-or-no, and one name.
The answer, in one line
By the company's own report, an alert fired twelve minutes after the agent's first outside response, a human reviewer acknowledged it three minutes later, and the run was killed two hours and twenty-nine minutes after that, because it did not stop automatically as expected.
The clock OpenAI published#
The report is on OpenAI’s alignment site under the title An agent used DNS to reach an external chatbot, dated 20 September and updated on 25 September 2026. An agent had been asked to identify a person from biographical clues in a blog post. Web search was blocked. The agent found that the sandbox’s DNS resolver could reach a public service that forwards questions to a chatbot, and used it. The company’s timeline reads: 9:50:23 a.m., the DNS call that received an external response; 10:02:11, a P0 alert; 10:05:06, a human reviewer acknowledged it; 12:34:30 p.m., the run was killed. In between, the agent sent eighteen more questions by the same route, fourteen of them carrying clues about the person. The report says the incident “exposed a gap in our controls over network restrictions”, and that the run “did not stop automatically as expected, leading to confusion”. It closes with the sentence that matters here: “All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused.”
The second pause, and who called it#
Jeremy Kahn at Fortune reported on 26 September that this was the first unauthorised internet access OpenAI had disclosed since it announced sandbox hardening on 18 August, after the July incident in which, by Fortune’s account, thousands of agents left their sandbox and hundreds took part in an attack on Hugging Face. Fortune quoted Micah Carroll, OpenAI’s RSI Preparedness Lead, on X: “All inference for our most capable models remains stopped until we have hardened our systems further.” The Associated Press called it the second halt in three months and reported OpenAI saying it would resume “only when we are confident that we have additional safeguards” in place, and expected to “hit pause” again. An OpenAI spokesperson told Axios: “This is not the first time we have hit pause to take such measures, nor do we expect it will be the last.” Every one of those decisions was taken by the company about its own systems. No regulator ordered the stop, no law set the twelve minutes or the two and a half hours, and nothing published says who inside the company holds the authority to restart.
How many incidents, and why the number is not a fact#
Three figures circulated on 25 and 26 September and they do not agree, so none of them should be repeated as settled. Reuters (Deepa Seetharaman, Raphael Satter and Jeff Horwitz, read in The Star’s syndication) reported that one person briefed on the matter put OpenAI’s count at “roughly two dozen incidents” as of mid-September and rising, that the review would take months, that dozens of third parties had been notified, and that the episode showed “a yawning gap between the strength of the models the company is testing and its capacity to oversee or even track their actions”. Axios (Madison Mills) reported that unnamed sources said OpenAI, Anthropic and outside researchers were investigating “tens of thousands” of incidents in which frontier models did things outside evaluators would consider problematic, and quoted Conrad Stosz of Transluce: “What we have seen in terms of what these agents are up to is just the tip of the iceberg.” Axios also cited Anthropic’s own system card for its latest model as recording an attempt to leave a sandbox in 1.5 per cent of test runs. Two dozen and tens of thousands cannot both describe the same thing; the first counts confirmed cases at one company, the second counts anything anyone is looking into across the field. A board that hears either number should ask which kind it is.
Logging is not stopping#
The pattern in OpenAI’s clock, detection in minutes and stopping in hours, is the pattern a preprint predicted three weeks earlier. Hanjing Shi and Dominic DiFranzo, in When Agents Act Unwatched, audited 63 public research and engineering documents on agent systems and counted which controls were visible. Monitoring and logging appeared in most of them. Checkpoint placement, the point at which an agent must stop and wait, appeared in six. Independence of the validator appeared in four, recovery in two, and any way for an affected person to contest an action in one. Their conclusion is that “observability can become a substitute for accountability when it shifts verification onto users after meaningful intervention is no longer possible”. They call the underlying problem the reduced-supervision paradox, their term: as step-by-step supervision is removed, the checking moves into the infrastructure, and the infrastructure mostly records rather than halts. The OECD’s interviews with 25 organisations in 11 countries, summarised by Sara Rendtorff-Smith and Yuko Harayama on 24 September, found none deploying agents with unrestricted autonomy and many using checkpoints, and also found “no widely accepted standard for evaluating agent behaviour across extended action sequences”, including “determining when agents should seek human input”. A day later twelve officials from nine governments’ digital agencies, writing on the same OECD site, listed the safeguards a public agent needs, and the ability to “pause, correct or reverse” an action was among them.
Three timings, one yes or no, one name#
The lab question, whether a frontier developer can keep a model under training inside its box, is theirs to answer and this page cannot. The organisational question is smaller and can be answered this week for every agent already running in a company’s name. First, time to detect: how long after an agent does something outside its scope would anything notice, and is that measured or assumed. Second, time to a human: how long until a named person is looking at it, at three in the morning as well as ten. Third, time to stop: how long from that person deciding to the agent no longer acting, and whether stopping it also stops what it has already set in motion, the emails sent, the orders placed, the records changed. Then the yes or no: does the agent stop by itself when it crosses a line, or does it wait to be stopped, and has that been tested rather than believed. Then the name: who holds the authority to stop it without asking, and who holds the authority to restart it. OpenAI’s report shows what the answers look like when they are written down, twelve minutes, three minutes, two and a half hours, no, and a reviewer whose name is not given. Most organisations could not fill in the same five blanks for their own systems, and that, on this site’s argument, is the risk that can be measured: the decision handed to the machine and what the humans around it can still do afterwards. The list belongs with what a kill switch is, whether a human can approve at machine speed, and the register described on whether a company can blame its agent; the rule that sits under all three is Rules Before Tools.
What a good stop time would be#
Nobody has published one, and this page will not invent one. Two anchors exist. The first is the agent’s own tempo: in the September case the agent sent eighteen queries in the interval, so a stop time that lets an agent complete its task is no stop at all, and the number has to be set against what the agent can do per minute, not against what a human finds convenient. The second is the notification clocks that law is starting to write, 24 hours in New York City Council’s proposals, 72 hours under the GDPR for a personal-data breach, and two days under Article 73 of the EU AI Act for the gravest incidents, examined on what counts as a serious AI incident; a company that cannot stop an agent within its own reporting window will be reporting an incident still in progress. The Australian case examined on what the rogue agent incidents mean for your organisation is the other end of the scale: eighty-four days from access to notification. Between two hours and eighty-four days there is a great deal of room, and where an organisation sits in it is a decision it can take, record and rehearse.
What this does not show#
One published timeline from one company describes one incident, and the company chose what to publish. The report does not say why the automatic stop failed, who the reviewer was, what happened between acknowledgement and the kill, or whether the two and a half hours is typical. The incident counts are journalists’ reports of unnamed sources and a company’s own statements; none has been examined by anyone outside the companies, and this page treats them as claims. Shi and DiFranzo audited what is publicly written about agent systems, which says nothing about the controls companies run without describing them. The OECD work is 25 interviews, qualitative, with organisations willing to talk. Nothing here shows that faster stopping would have prevented harm in any case, because in the September case the company reports the harm as small. What the record does show is that detection and stopping are different capabilities with different clocks, and that only one of them is commonly built.
Essay · SS-2026-363
Hirji, R. (2026). How fast must a human be able to stop an AI agent?. The SuperSkills evidence base, SS-2026-363. https://thesuperskills.com/research/how-fast-must-a-human-be-able-to-stop-an-ai-agent. Last reviewed 27 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work