← Research
Research

How do you audit an AI-assisted decision?

Most organisations can produce the system log and cannot produce the human decision. The second is what regulators have been asking about.

Last reviewed: 27 August 2026

What Article 12 and Article 14 actually require, what the Court of Justice has held explanation to mean, the two cases where the record did not exist, and the six questions a real audit trail answers.

Questions this page answersAll 780 questions this research covers

By reconstructing who decided what, on what basis, with what authority to refuse, and being able to show it afterwards to someone who was not there. Most organisations can produce the system log and cannot produce the human decision, which is the half that regulators and courts have actually been asking about.

This page separates what the law now requires from what has become industry habit, because the two are diverging and the habit is the more confident of the two.

What the law actually requires#

Logging is now a legal obligation. Article 12 of the EU AI Act requires that high-risk AI systems technically allow the automatic recording of events over the system's lifetime, so that risk situations can be identified, post-market monitoring can happen, and operation can be monitored. Article 19 requires providers to keep those logs for at least six months.

Note the shape of that obligation. It requires the capability to reconstruct what the system did. It does not require anyone to read the logs, and it says nothing about recording what the human was thinking.

Human oversight is a capability requirement, not a review requirement. Article 14 requires that overseers can understand the system's capacities and limits, remain aware of the tendency to over-rely on output, interpret output correctly, decide not to use the system, override or reverse its output, and interrupt it.

Here is the part almost everyone gets wrong. The AI Act does not require a human to review every high-risk AI decision. The only per-decision requirement is in Article 14(5), and it applies to remote biometric identification: no action may be taken on an identification unless it has been separately verified by at least two natural persons, with a carve-out for certain law enforcement and border uses. Everywhere else the duty is that oversight is possible and capable, not that it happens case by case.

Which means an organisation that has installed per-decision sign-off believing it is a compliance requirement has bought the configuration the evidence says performs worst, for a reason that is not in the regulation.

In February 2025 the Court of Justice of the European Union decided Dun and Bradstreet Austria, on what GDPR means by meaningful information about the logic involved in an automated decision.

The Court held that a controller must describe the procedure and principles actually applied, in a way that lets the person understand which of their data was used and how. It may be appropriate to show the extent to which different data would have produced a different result. And two limits that matter operationally: disclosing the algorithm is not a sufficient explanation, and a blanket refusal on trade-secret grounds is not permitted, with contested material going to the supervisory authority or court to balance instead.

The standard is therefore comprehension by the affected person, not disclosure to them. Model cards and technical documentation do not discharge it. Nor does source code, which the ruling does not require.

Two cases where the record did not exist#

The instructive enforcement so far is not about models behaving badly. It is about organisations unable to evidence a human.

In 2023 the Amsterdam Court of Appeal considered drivers deactivated after fraud flags, and found the human intervention in the automated decisions amounted to little more than a symbolic act. The company could not show that a real decision-maker, informed by the specifics, had made the call. The court also rejected a trade-secret defence for withholding algorithmic information.

In 2021 the Italian data protection authority fined Foodinho, the Glovo subsidiary, 2.6 million euros, finding it had not sufficiently explained how its automated order management worked, had not ensured the accuracy of algorithmic performance scoring, and had given riders no route to challenge an algorithmic decision or obtain human intervention. A second and larger fine followed.

In both, the system was not the finding. The absence of a person who could be shown to have decided was the finding.

The frameworks, and what they are not#

Three things get cited as if they were compliance. None of them is.

NIST AI Risk Management Framework 1.0, published January 2023, is voluntary. It gives a common vocabulary organised as Govern, Map, Measure, Manage, with a playbook and a generative AI profile. Its value is that a board will recognise the structure. It confers no legal status.

ISO/IEC 42001:2023 is a certifiable management system standard, published December 2023, built on Plan-Do-Check-Act. An organisation can be certified against it by a third party. Certification demonstrates that a management system exists and is operating. It is not equivalent to meeting the EU AI Act, and the two are routinely conflated in procurement.

Guidance from national data protection authorities sits in the same category: useful, non-statutory, and not a defence on its own.

What an audit should actually be able to reconstruct#

Six questions. If the record cannot answer them, the audit trail is a system log rather than a decision record.

The first five are recordable at the moment of the decision at almost no cost. The sixth is a report anyone can run today.

Nobody has shown that any of this works#

Nobody has shown that audit trails change behaviour. This research could find no study, peer-reviewed or otherwise, testing whether the existence of an audit trail alters what decision-makers do, or improves outcomes, as against producing a record after the fact.

That is an uncomfortable thing for a page recommending audit trails to say. It remains the position. The case for them here rests on legal obligation and on the demonstrated cost of not having one, in the Uber and Foodinho decisions, rather than on evidence that they work. If someone publishes a study, this page changes.

There is a related caution. An audit trail that records approvals and not disagreements is the paperwork equivalent of usage theatre: it produces the evidence of oversight while measuring none of it.

On the legal duty, meaningful human oversight and what AI literacy means for leaders. On why review is the weakest position, human in the loop is not a safeguard and human and AI decision making. On the practical tool, the Delegation Boundary Map. On who carries the duty, who owns verification and who supervises work they cannot do. On the board version, what a board should ask.

Key research and primary sources

About this research#

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. Regulation is quoted from the published text of Regulation (EU) 2024/1689 and the Court's own press release. The Amsterdam and Garante decisions are summarised from convergent legal reporting rather than from the primary judgments, which is a weaker basis and is flagged here for that reason. That audit trails change behaviour is unevidenced, which the page states. Not legal advice.

How this research works  ·  Reviewed quarterly  ·  Found an error? Tell me and it is corrected on the page.

Cite this

Hirji, R. (2026). How do you audit an AI-assisted decision? The SuperSkills Intelligence Company. Last reviewed 27 August 2026. thesuperskills.com/research/how-do-you-audit-an-ai-assisted-decision

Questions answered on this page

How do you audit an AI-assisted decision?

By being able to reconstruct six things: who decided by name, what they were looking at, how long they had per item, whether they could have detected an error, on what stated grounds they could have disagreed, and whether anyone ever did. Most organisations can produce the system log and not the human decision. The enforcement so far, in the Amsterdam Court of Appeal ruling against Uber in 2023 and the Italian Garante's Foodinho decisions, has turned on the absence of an evidenced human decision-maker rather than on the model.

Does the EU AI Act require a human to review every AI decision?

No, and this is widely misunderstood. Article 14 requires that human oversight be possible and capable: overseers must be able to understand the system, remain aware of automation bias, interpret output correctly, decline to use it, override it and interrupt it. The only per-decision requirement is Article 14(5), which applies to remote biometric identification and requires separate verification by at least two natural persons, with carve-outs for certain law enforcement and border uses. For other high-risk systems the duty is capability, not case-by-case review.

What does the EU AI Act require organisations to log?

Article 12 requires that high-risk AI systems technically allow automatic recording of events over the system lifetime, to enable identification of risk situations, post-market monitoring and monitoring of operation. For remote biometric identification the logs must record the start and end time of each use, the reference database, the input data that produced a match, and who verified the result. Article 19 requires providers to retain the logs for at least six months. The obligation is to have the capability, not to read the logs.

Is ISO 42001 certification the same as complying with the AI Act?

No. ISO/IEC 42001:2023 is a certifiable management system standard: it demonstrates that an AI management system exists and operates on a Plan-Do-Check-Act basis. The NIST AI Risk Management Framework is voluntary and confers no legal status at all. Neither is equivalent to meeting the requirements of Regulation (EU) 2024/1689, and the two are routinely conflated in procurement.

In this hub

Judgement, oversight and accountability

Who decides, who checks, and who is answerable when the machine was involved.

Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Could you reconstruct how a decision was reached six months from now, to somebody hostile? Building a trail that holds up when it matters is the job. AI advisory for CEOs and boards.

Oversight is the topic most often agreed with in principle and least often implemented. There is the human oversight version, and the full range of topics and audiences.

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.