By reconstructing who decided what, on what basis, with what authority to refuse, and being able to show it afterwards to someone who was not there. Most organisations can produce the system log and cannot produce the human decision, which is the half that regulators and courts have actually been asking about.
This page separates what the law now requires from what has become industry habit, because the two are diverging and the habit is the more confident of the two.
What the law actually requires
Logging is now a legal obligation. Article 12 of the EU AI Act requires that high-risk AI systems technically allow the automatic recording of events over the system's lifetime, so that risk situations can be identified, post-market monitoring can happen, and operation can be monitored. Article 19 requires providers to keep those logs for at least six months.
Note the shape of that obligation. It requires the capability to reconstruct what the system did. It does not require anyone to read the logs, and it says nothing about recording what the human was thinking.
Human oversight is a capability requirement, not a review requirement. Article 14 requires that overseers can understand the system's capacities and limits, remain aware of the tendency to over-rely on output, interpret output correctly, decide not to use the system, override or reverse its output, and interrupt it.
Here is the part almost everyone gets wrong. The AI Act does not require a human to review every high-risk AI decision. The only per-decision requirement is in Article 14(5), and it applies to remote biometric identification: no action may be taken on an identification unless it has been separately verified by at least two natural persons, with a carve-out for certain law enforcement and border uses. Everywhere else the duty is that oversight is possible and capable, not that it happens case by case.
Which means an organisation that has installed per-decision sign-off believing it is a compliance requirement has bought the configuration the evidence says performs worst, for a reason that is not in the regulation.
Explanation now has a legal standard, and it is comprehension
In February 2025 the Court of Justice of the European Union decided Dun and Bradstreet Austria, on what GDPR means by meaningful information about the logic involved in an automated decision.
The Court held that a controller must describe the procedure and principles actually applied, in a way that lets the person understand which of their data was used and how. It may be appropriate to show the extent to which different data would have produced a different result. And two limits that matter operationally: disclosing the algorithm is not a sufficient explanation, and a blanket refusal on trade-secret grounds is not permitted, with contested material going to the supervisory authority or court to balance instead.
The standard is therefore comprehension by the affected person, not disclosure to them. Model cards and technical documentation do not discharge it. Nor does source code, which the ruling does not require.
Two cases where the record did not exist
The instructive enforcement so far is not about models behaving badly. It is about organisations unable to evidence a human.
In 2023 the Amsterdam Court of Appeal considered drivers deactivated after fraud flags, and found the human intervention in the automated decisions amounted to little more than a symbolic act. The company could not show that a real decision-maker, informed by the specifics, had made the call. The court also rejected a trade-secret defence for withholding algorithmic information.
In 2021 the Italian data protection authority fined Foodinho, the Glovo subsidiary, 2.6 million euros, finding it had not sufficiently explained how its automated order management worked, had not ensured the accuracy of algorithmic performance scoring, and had given riders no route to challenge an algorithmic decision or obtain human intervention. A second and larger fine followed.
In both, the system was not the finding. The absence of a person who could be shown to have decided was the finding.
The frameworks, and what they are not
Three things get cited as if they were compliance. None of them is.
NIST AI Risk Management Framework 1.0, published January 2023, is voluntary. It gives a common vocabulary organised as Govern, Map, Measure, Manage, with a playbook and a generative AI profile. Its value is that a board will recognise the structure. It confers no legal status.
ISO/IEC 42001:2023 is a certifiable management system standard, published December 2023, built on Plan-Do-Check-Act. An organisation can be certified against it by a third party. Certification demonstrates that a management system exists and is operating. It is not equivalent to meeting the EU AI Act, and the two are routinely conflated in procurement.
Guidance from national data protection authorities sits in the same category: useful, non-statutory, and not a defence on its own.
What an audit should actually be able to reconstruct
Six questions. If the record cannot answer them, the audit trail is a system log rather than a decision record.
- Who decided, by name. Not the team, not the function. Attribution assigned after an outcome is not accountability.
- What were they looking at. The output, and what else. If they saw only the recommendation, the record should say so.
- How long did they have. Time per item is the single most diagnostic number in any oversight arrangement, and almost nobody records it.
- Could they have detected an error. The capability test. Could this person have produced the work themselves, well enough to notice if it were wrong? If not, the control is nominal.
- On what grounds could they have disagreed. Stated in advance. A reviewer with no defined basis for objecting will not object.
- Did anyone ever disagree. An override rate of zero across a quarter evidences an untested right rather than a good system.
The first five are recordable at the moment of the decision at almost no cost. The sixth is a report anyone can run today.
Nobody has shown that any of this works
Nobody has shown that audit trails change behaviour. This research could find no study, peer-reviewed or otherwise, testing whether the existence of an audit trail alters what decision-makers do, or improves outcomes, as against producing a record after the fact.
That is an uncomfortable thing for a page recommending audit trails to say, and it is the position. The case for them here rests on legal obligation and on the demonstrated cost of not having one, in the Uber and Foodinho decisions, rather than on evidence that they work. If someone publishes a study, this page changes.
There is a related caution. An audit trail that records approvals and not disagreements is the paperwork equivalent of usage theatre: it produces the evidence of oversight while measuring none of it.
Related SuperSkills research
On the legal duty, meaningful human oversight and what AI literacy means for leaders. On why review is the weakest position, human in the loop is not a safeguard and human and AI decision making. On the practical tool, the Delegation Boundary Map. On who carries the duty, who owns verification and who supervises work they cannot do. On the board version, what a board should ask.
Key research and primary sources
- Article 12, Record-keeping, Regulation (EU) 2024/1689, and Article 14, Human oversight.
- Court of Justice of the European Union (2025). Dun and Bradstreet Austria, Case C-203/22, judgment of 27 February 2025.
- National Institute of Standards and Technology (2023). AI Risk Management Framework 1.0.
- International Organization for Standardization (2023). ISO/IEC 42001:2023, Artificial intelligence management system.
- Vaccaro, M., Almaatouq, A. and Malone, T. (2024). When combinations of humans and AI are useful. Nature Human Behaviour, 8.
About this research
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026) and founder of The SuperSkills Intelligence Company. Regulation is quoted from the published text of Regulation (EU) 2024/1689 and the Court's own press release. The Amsterdam and Garante decisions are summarised from convergent legal reporting rather than from the primary judgments, which is a weaker basis and is flagged here for that reason. That audit trails change behaviour is unevidenced, which the page states. Not legal advice. Reviewed quarterly.
Cite this
Hirji, R. (2026). How do you audit an AI-assisted decision? The SuperSkills Intelligence Company. Last reviewed 27 August 2026. thesuperskills.com/research/how-do-you-audit-an-ai-assisted-decision