← Research
Research

How do you audit an AI-assisted decision?

Most organisations can produce the system log and cannot produce the human decision. The second is what regulators have been asking about.

Last reviewed: 27 August 2026

What Article 12 and Article 14 actually require, what the Court of Justice has held explanation to mean, the two cases where the record did not exist, and the six questions a real audit trail answers.

By reconstructing who decided what, on what basis, with what authority to refuse, and being able to show it afterwards to someone who was not there. Most organisations can produce the system log and cannot produce the human decision, which is the half that regulators and courts have actually been asking about.

This page separates what the law now requires from what has become industry habit, because the two are diverging and the habit is the more confident of the two.

What the law actually requires

Logging is now a legal obligation. Article 12 of the EU AI Act requires that high-risk AI systems technically allow the automatic recording of events over the system's lifetime, so that risk situations can be identified, post-market monitoring can happen, and operation can be monitored. Article 19 requires providers to keep those logs for at least six months.

Note the shape of that obligation. It requires the capability to reconstruct what the system did. It does not require anyone to read the logs, and it says nothing about recording what the human was thinking.

Human oversight is a capability requirement, not a review requirement. Article 14 requires that overseers can understand the system's capacities and limits, remain aware of the tendency to over-rely on output, interpret output correctly, decide not to use the system, override or reverse its output, and interrupt it.

Here is the part almost everyone gets wrong. The AI Act does not require a human to review every high-risk AI decision. The only per-decision requirement is in Article 14(5), and it applies to remote biometric identification: no action may be taken on an identification unless it has been separately verified by at least two natural persons, with a carve-out for certain law enforcement and border uses. Everywhere else the duty is that oversight is possible and capable, not that it happens case by case.

Which means an organisation that has installed per-decision sign-off believing it is a compliance requirement has bought the configuration the evidence says performs worst, for a reason that is not in the regulation.

Explanation now has a legal standard, and it is comprehension

In February 2025 the Court of Justice of the European Union decided Dun and Bradstreet Austria, on what GDPR means by meaningful information about the logic involved in an automated decision.

The Court held that a controller must describe the procedure and principles actually applied, in a way that lets the person understand which of their data was used and how. It may be appropriate to show the extent to which different data would have produced a different result. And two limits that matter operationally: disclosing the algorithm is not a sufficient explanation, and a blanket refusal on trade-secret grounds is not permitted, with contested material going to the supervisory authority or court to balance instead.

The standard is therefore comprehension by the affected person, not disclosure to them. Model cards and technical documentation do not discharge it. Nor does source code, which the ruling does not require.

Two cases where the record did not exist

The instructive enforcement so far is not about models behaving badly. It is about organisations unable to evidence a human.

In 2023 the Amsterdam Court of Appeal considered drivers deactivated after fraud flags, and found the human intervention in the automated decisions amounted to little more than a symbolic act. The company could not show that a real decision-maker, informed by the specifics, had made the call. The court also rejected a trade-secret defence for withholding algorithmic information.

In 2021 the Italian data protection authority fined Foodinho, the Glovo subsidiary, 2.6 million euros, finding it had not sufficiently explained how its automated order management worked, had not ensured the accuracy of algorithmic performance scoring, and had given riders no route to challenge an algorithmic decision or obtain human intervention. A second and larger fine followed.

In both, the system was not the finding. The absence of a person who could be shown to have decided was the finding.

The frameworks, and what they are not

Three things get cited as if they were compliance. None of them is.

NIST AI Risk Management Framework 1.0, published January 2023, is voluntary. It gives a common vocabulary organised as Govern, Map, Measure, Manage, with a playbook and a generative AI profile. Its value is that a board will recognise the structure. It confers no legal status.

ISO/IEC 42001:2023 is a certifiable management system standard, published December 2023, built on Plan-Do-Check-Act. An organisation can be certified against it by a third party. Certification demonstrates that a management system exists and is operating. It is not equivalent to meeting the EU AI Act, and the two are routinely conflated in procurement.

Guidance from national data protection authorities sits in the same category: useful, non-statutory, and not a defence on its own.

What an audit should actually be able to reconstruct

Six questions. If the record cannot answer them, the audit trail is a system log rather than a decision record.

The first five are recordable at the moment of the decision at almost no cost. The sixth is a report anyone can run today.

Nobody has shown that any of this works

Nobody has shown that audit trails change behaviour. This research could find no study, peer-reviewed or otherwise, testing whether the existence of an audit trail alters what decision-makers do, or improves outcomes, as against producing a record after the fact.

That is an uncomfortable thing for a page recommending audit trails to say, and it is the position. The case for them here rests on legal obligation and on the demonstrated cost of not having one, in the Uber and Foodinho decisions, rather than on evidence that they work. If someone publishes a study, this page changes.

There is a related caution. An audit trail that records approvals and not disagreements is the paperwork equivalent of usage theatre: it produces the evidence of oversight while measuring none of it.

Related SuperSkills research

On the legal duty, meaningful human oversight and what AI literacy means for leaders. On why review is the weakest position, human in the loop is not a safeguard and human and AI decision making. On the practical tool, the Delegation Boundary Map. On who carries the duty, who owns verification and who supervises work they cannot do. On the board version, what a board should ask.

Key research and primary sources

About this research

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026) and founder of The SuperSkills Intelligence Company. Regulation is quoted from the published text of Regulation (EU) 2024/1689 and the Court's own press release. The Amsterdam and Garante decisions are summarised from convergent legal reporting rather than from the primary judgments, which is a weaker basis and is flagged here for that reason. That audit trails change behaviour is unevidenced, which the page states. Not legal advice. Reviewed quarterly.

Cite this

Hirji, R. (2026). How do you audit an AI-assisted decision? The SuperSkills Intelligence Company. Last reviewed 27 August 2026. thesuperskills.com/research/how-do-you-audit-an-ai-assisted-decision

In this hub

AI and Human Judgement

Does AI weaken judgement? The evidence, and what to do about it.

The work

Where the writing comes from.

These essays draw on research across more than 200 organisations in 30 countries. See the wider body of work, or bring it into your organisation.

All research →