- Should leaders use AI themselves?
- What should a chief executive understand personally before approving AI deployment?
- Who decides which human capabilities are worth preserving?
- What is our appetite for AI risk?
- Which AI decisions should come to the board?
- What should management report to the board about AI every quarter?
- What would an effective board AI dashboard contain?
- How do we know our AI controls actually work?
- How should directors themselves use AI?
Most boards now have an AI item. Fewer have a clear picture of what oversight of it actually consists of, beyond a quarterly paper and a risk register line. This page sets out what the frameworks require, what belongs on an agenda, and the one question none of the frameworks answers.
What the frameworks actually require#
Three documents do most of the work in board conversations, and they are worth knowing at the level of what they oblige rather than what they signal.
- NIST AI Risk Management Framework 1.0 organises around four functions, Govern, Map, Measure and Manage, with a Playbook and a 2024 generative AI profile. It gives a board a shared vocabulary it will recognise. It is voluntary and confers no legal status, so adopting it is a decision rather than a compliance step.
- Article 14 of the EU AI Act is the sharpest text on oversight anywhere. High-risk systems must be designed so natural persons can effectively oversee them, and those persons must be enabled to understand the system’s capacities and limits, to remain aware of the tendency to over-rely on output, with automation bias named in the legislation itself, to interpret output correctly, and to decide not to use it, disregard it, override it, reverse it, or stop it. Note what that treats as the content of oversight: the capability to override, not the presence of a person. The provisions do not apply until December 2027 at the earliest.
- Article 12, with Article 19, requires high-risk systems to allow automatic recording of events across the system lifetime, and providers to keep those logs for at least six months. The ability to reconstruct what a system did is now a legal requirement rather than good practice. Nothing in it requires anybody to read the logs.
The vocabulary, and where it stops working#
Boards run oversight through structures that predate this technology and mostly transfer well. Three lines of defence puts operational management first, risk and compliance second, internal audit third. Assurance mapping asks who is testing what and where the duplication and the holes are. Risk appetite sets how much of a given exposure the board is willing to carry. Every one of those applies to AI, and a board that already runs them has most of the machinery.
The difficulty is specific rather than general. The capability question does not sit in any of the three lines. The first line runs the system. The second checks that policy was followed. The third audits whether the second did its job. None of them is asked whether the organisation’s people could still do the work if the system stopped, and none of them would be at fault for not asking. That is the gap this research exists to describe, and it sits in the assurance map rather than in anybody’s performance.
Why a human in the loop is not a control#
The most common answer to a board question about AI risk is that a person reviews the output. Vaccaro, Almaatouq and Malone’s preregistered meta-analysis of 106 experimental studies and 370 effect sizes found human and AI combinations performed significantly worse on average than the better of human or AI alone, at Hedges’ g of -0.23, with the losses concentrated in decision-making.
Their conclusion belongs in front of any board being offered a human reviewer as a mitigation: adding a human is not a control, and undesigned pairing can subtract. Article 14 arrives at the same place from the legal side by specifying competence, training and authority together rather than presence. A named reviewer without the standing to say no is a control on paper. See human in the loop is not a safeguard for the longer version.
Reversal has to exist before it is needed#
The NIST Playbook, at MANAGE 2.4, requires mechanisms and assigned responsibilities to supersede, disengage or deactivate systems performing inconsistently with intended use, and names five triggering conditions: end of system lifetime; risks exceeding tolerance thresholds; mitigation beyond the organisation’s capacity; feasible mitigations failing regulatory, legal or normative standards; and impending risk detected in monitoring.
An authoritative framework therefore treats deployment as reversible and expects the mechanism, its thresholds and its fallback to be in place beforehand. The board question follows directly: who holds the authority to stop this, what would trigger it, and what happens to the work in the meantime. Thresholds set while everyone is pleased with a system are the only ones worth having, because the people defending a decision cannot set them afterwards.
Accountability is arriving through the courts, and it runs upward#
Directors should know about Ayinde v London Borough of Haringey and Al-Haroun v Qatar National Bank, heard together in 2025 under the Hamid jurisdiction after fabricated citations were placed before the court. In Ayinde, five cited authorities did not exist.
The court held at [6] that freely available generative AI tools are not capable of conducting reliable legal research, and at [7] that those using them carry a professional duty to check accuracy against authoritative sources. Two further points matter more for governance than the headline. At [8] the duty extends to lawyers relying on other people’s AI-assisted work. At [81] a lawyer is not entitled to rely on their client for accuracy.
In England and Wales the verification duty is therefore settled, non-delegable, and travels upward to whoever supervises. This is a professional obligations case rather than a company law one, and no reported case yet covers somebody who checked competently and was misled anyway. The direction is clear enough for a board to act on: accountability for AI-assisted output does not stop at the person who produced it.
One national framework has already written in the capability risk#
Boards told that capability loss is a soft concern should see Singapore’s Model AI Governance Framework for Agentic AI, version 1.0, at section 2.4.3.
As agents take over entry level tasks, which typically serve as the training ground for new staff, this could lead to loss of basic operational knowledge for the users. Organisations should identify core capabilities of each job and provide sufficient training and work exposure so that users retain foundational skills.
IMDA, Model AI Governance Framework for Agentic AI, version 1.0, section 2.4.3
A national government has put the removal of entry-level work, and the loss of the training ground that goes with it, into an operative governance framework. It is guidance rather than law, it states a risk and a duty to train rather than evidence that deskilling has occurred, and it sets no measurement. What it establishes is that the missing rungs argument is now inside the governance literature rather than outside it, which changes what a director can reasonably say they had not considered.
What actually goes on the agenda#
The practical shape, for a board meeting quarterly. None of this requires technical depth, and all of it requires somebody to have prepared an answer.
- The declared position, once a year. Whether AI is being used to augment or to replace, by domain, in writing. If the business cases count headcount as the benefit while the paper says augmentation, the board is reading the wrong document. See who should own AI strategy.
- The capability floor, with an owner and a date. What the organisation must still be able to do unaided in three years, who reports on whether it still can, and when they last checked. This is the item nobody currently owns.
- Decision rights, in writing, before an incident. Who decides what, who can override a system, and whether anybody has actually done so recently. A recorded override is evidence the function is real. See who can override an AI system.
- Reversal thresholds, tested rather than rewritten. Per MANAGE 2.4, with the fallback named.
- One reconstructed decision per quarter. Take a real AI-assisted decision and ask the organisation to show how it was reached. Article 12 logging makes this possible; nothing makes it happen. See how to audit an AI-assisted decision.
On what a paper should contain, the test is simple. If the reporting is adoption rates, licences issued or hours saved, the board is being shown procurement. The question a paper should answer is whether anybody got better at anything, and whether anybody could still do the work without the system. For the full set of questions, see what a board should ask about AI.
Where this sits in my own argument#
My position throughout this research is that a governance framework designs the control, and the work I do tests whether the human part of it operates. The frameworks above are good. Read them closely and you find that each one specifies what should exist rather than whether it functions, which is reasonable, since that is what frameworks are for.
The oversight failure I keep meeting involves no absent control at all. Something exists, is documented, has a named owner, and would not catch anything. That is drift in its governance form: nobody decided the oversight would be decorative, and nobody has checked.
What I have observed in organisations#
The clearest thing I have seen about board-level oversight is how much of it depends on whether the senior team uses the tools at all. In one accounting firm, several partners were slow to use ChatGPT or Copilot for anything, including email. The question of what AI meant for the firm was pushed down the agenda and treated as a process matter, because the people at the top had no working feel for what they were governing.
It changed when the chair of the board changed the ethos around AI. Not a new policy, and not a framework. A different expectation from the top about whether this was something the board engaged with personally.
Which answers a question people ask me often. Should leaders use these tools themselves? Yes, and not for productivity. A director who has never watched a model produce a confident, wrong answer in their own domain has no calibration for the risk they are being asked to oversee, and no way to tell a real control from a described one.
What this page does not claim#
It does not claim the frameworks are inadequate. NIST gives a board vocabulary it can use, Article 14 is the most serious text written on human oversight, and Article 12 makes reconstruction a legal requirement. Each is doing its job.
It does not claim any of this is happening. The Article 14 oversight provisions do not apply until December 2027 at the earliest, NIST is voluntary, the Singapore framework is guidance, and nothing in Article 12 requires anybody to read a log. What exists is an obligation to be able to; whether organisations do is unmeasured.
And it offers no legal advice. Ayinde is a professional obligations case in England and Wales, and how the duty to verify translates into directors’ duties in a given jurisdiction is a question for counsel rather than for this page.
Cite this
Hirji, R. (2026). What board oversight of AI actually looks like. The SuperSkills Intelligence Company. Last reviewed 1 September 2026. thesuperskills.com/research/what-board-oversight-of-ai-looks-like