Discussion of agents inside organisations tends to be about capability: what the system can do, how reliably, on what sort of task. The question that decides the exposure is duller and more specific. What can it read, what can it change, what can it start, and who can stop it once it has started.
Those four are permissions rather than capabilities. They are decided by whoever configures the system, usually early, usually in a hurry, and almost never reviewed.
A regulator has recorded the first case#
The Spanish data protection authority published, in September 2026, its account of the first personal data breach notified to it that was carried out by an agent. On the regulator's description the agent chained the phases of the attack itself: it searched for weaknesses, logged in, found application vulnerabilities, modified personal data and reached invoices. The authority describes this as a shift from AI-assisted attack to autonomous action.
It is one case in one jurisdiction, reported from a breach notification rather than established by investigation, and the post does not say how much human direction was involved or whether the notifying organisation's account is accurate. Taken at that weight it still marks the point where the question stopped being hypothetical for a European regulator.
Nobody can reliably say what is running#
Chan and colleagues set out what visibility into deployed agents would require, defined as information about where, why, how and by whom they are used, and assessed three families of measure: agent identifiers, real-time monitoring and activity logging. They name five agent-specific risks, and the last of them is the one that undermines any inventory: sub-agents, where an agent starts further agents and the record of what is running stops being complete.
The authors present options for further study rather than recommendations, and offer no evidence that any measure works. The negative finding is the usable one. Knowing what is running, and what it has started, is an open technical problem. An organisation that has written a policy requiring an inventory of agents has written a policy requiring something nobody can currently supply.
Published practice describes watching, not stopping#
Shi and DiFranzo audited 63 public artefacts on agent systems, 46 research papers and 17 engineering, documentation, security and governance sources, counting which accountability controls were described.
Monitoring and tool mediation appeared in about 40 of them. Checkpoint placement appeared in 6. Validator independence in 4. Recovery in 2. Contestability in 1. Their conclusion is that observability can become a substitute for accountability when it shifts verification onto users after meaningful intervention is no longer possible.
This is a preprint, the abstract rather than the full text was read for the graded entry, and it counts what is publicly written rather than what organisations actually run. As a measure of the published state of the art it is sobering: the field has largely solved looking and has barely started on stopping, undoing and appealing.
Two instruments worth knowing about#
Dubai International Financial Centre Regulation 10 states that human-defined processing purposes must always prevail in the development and use of autonomous systems, and draws the analogy directly: where a system operates for the benefit of its deployer, its position is substantially similar to that of an employee of the deploying organisation. It is confined to one free zone, addresses liability rather than competence, and requires nothing about the capability of the responsible person.
In the other direction, the United States banking agencies replaced their model risk framework with SR 26-2 in April 2026 and narrowed the definition of a model to a complex quantitative method applying statistical, economic or financial theories to produce quantitative estimates, expressly placing generative AI outside scope. The same footnote directs firms to their own risk management and governance practices for tools outside it. Anybody citing model risk management as the control covering a generative system should read the new scope first.
The four questions, asked properly#
- What can it read? Not what it uses. What it could open if the task drifted, including the mailbox, the drive and the systems reached through a connector somebody approved once.
- What can it change? Reading is recoverable. Writing, sending, posting, deleting and paying are the verbs worth separating out, and they are usually bundled with the rest.
- What can it start? A system that can launch another system makes its own inventory incomplete, and the inventory is the basis of every other control.
- Who can stop it, from where, at three in the morning? A named person, a route they have actually used, and a test that has actually been run.
A fifth question matters as much and is asked least often: what does a person do after it has gone wrong. Recovery appeared in 2 of 63 artefacts, which is the measure of how little attention that half has had.
What would settle it#
A survey of what organisations have configured rather than what they describe: permissions actually granted, stop routes actually tested, recovery actually exercised. The audit above counts public artefacts and says so. Nothing in this base measures deployed practice, so the gap between the controls firms describe and the controls they run is unmeasured in both directions.
Where this sits in my own argument#
Rules before tools is the general form of this: the configuration decision is a governance decision, taken before anybody can see the consequences. An agent with wide reach and no tested stop is drift in its most literal sense, because the organisation has not decided anything, it has accepted a default.
Related SuperSkills research#
On accountability for an agent's actions, can a company blame its AI agent? On approval at machine speed, can a human approve an AI decision at machine speed? On oversight that cannot function, human in the loop is not a safeguard and meaningful human oversight.
Key sources
- Agencia Española de Protección de Datos (2026). First notified breach executed by an AI agent.
- Chan, A. et al. (2024). Visibility into AI Agents.
- Shi, H. and DiFranzo, D. (2026). When Agents Act Unwatched: The Reduced-Supervision Paradox in Agentic AI.
- DIFC Commissioner of Data Protection (2024). Regulation 10 on autonomous and semi-autonomous systems.
- Board of Governors of the Federal Reserve System et al. (2026). Supervisory Guidance on Model Risk Management, SR 26-2.
About this research#
Written by Rahim Hirji, author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company.
How this research works · Reviewed quarterly · Found an error? Tell me and it is corrected on the page.
Evidence review · SS-2026-394 · Graded against the published rubric · 1 peer-reviewed study, 1 working paper, 1 statutory investigation and 2 of other kinds
Hirji, R. (2026). What Your Agents Can Reach. The SuperSkills evidence base, SS-2026-394. https://thesuperskills.com/research/what-your-agents-can-reach. Last reviewed 3 October 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work