- How does the board know management's claims about AI are true?
- Does the board need independent assurance over AI?
Almost every AI figure a board is shown was produced by the people whose programme it measures. Time saved, licences activated, adoption, productivity: self-reported, by the enthusiasts, without a baseline. That is no accusation; the best-conducted trials say it about themselves. The UK Government Digital Service's cross-government Copilot experiment reported 26 minutes a day saved and stated that it could not identify how the saved time was spent; the Department for Work and Pensions' evaluation reported 19 minutes and listed, in its own limitations chapter, the absence of baseline data, self-selection towards enthusiasts and acquiescence bias on the time question. If the civil service says that about its own numbers, a board should assume it of a vendor deck.
The answer, in one line
Mostly it does not, because the claims are self-reported by the people whose programme they describe. The UK government's own Copilot trials reported 26 and 19 minutes a day saved and said in their limitations chapters that there was no control group, no baseline and self-selection of enthusiasts.
Why self-report is the norm#
Measuring what AI did to output quality or decision quality is expensive and slow, and measuring what people say it did is neither. McKinsey's 2025 survey of 1,491 organisations, the largest of its kind, measures self-reported EBIT impact and finds its 25 attributes explain a fifth of the variance. MIT NANDA's 95 per cent rests on 52 interviews. Gartner's 40 per cent is a forecast. Every number in circulation about enterprise AI is a claim by an interested party, and management's numbers are drawn from the same well. The board is not being lied to. It is being shown the only kind of figure that exists.
What assurance would check#
Independent assurance over AI, where AI takes part in consequential decisions, checks four documents rather than a dashboard. The allocation: which decisions the machine may inform, recommend or execute, which it may never own. The owners: a name against each. The capability floor: what the organisation keeps doing unaided so that somebody can still tell when the machine is wrong, and the evidence that they still can. And the disagreement rate: how often the humans reviewing machine output reach a different answer, measured rather than assumed. Vaccaro and colleagues' meta-analysis of 106 experiments found human and AI combinations underperforming the better of the two alone, with the losses in decision-making; a review that never disagrees is the mechanism by which that happens. A programme that cannot report the rate has no oversight, only a signature. One that reports zero has a rubber stamp.
The strongest argument for it came from the builders#
In September 2026 the chief executive of Anthropic asked, in public, for outside evaluators embedded inside AI companies with employee-level access, on the grounds that his own company's assurances about its own systems were not enough for anybody to rely on. Microsoft published rules for its models and opened them to six weeks of public consultation rather than announcing them as settled. Both are arguments, from the people with the most to lose by making them, that self-report is insufficient at the frontier. A board being told that its own management's self-report is sufficient is being asked to hold a lower standard than the laboratories now hold themselves to.
What nobody has measured#
There is no study of AI assurance and its effect on outcomes, and the four-document review described here is a proposal drawn from the model risk regime in banking rather than a tested practice. The evidence on this page is that the figures boards receive are self-reported, which their own producers confirm, and that the failures of AI programmes are consistently located in unmade decisions, which is what the four documents record. Whether reviewing them independently prevents failure is untested. That it makes the failure visible earlier is the claim.
What the board asks for#
Not a better dashboard. The four documents, once, and then whenever a system's scope changes; an annual external reading of them by somebody who does not report to the programme; and the disagreement rate every quarter, with the admission if it is not measured. The full list of questions is at what should a board ask about AI; what the audit trail has to contain is at how do you audit an AI-assisted decision; and the reason a governance pack does not answer this on its own is at AI governance versus AI leadership.
Key sources
- Government Digital Service (2025). Microsoft 365 Copilot Experiment: Cross-Government Findings Report. Graded entry.
- Department for Work and Pensions (2026). An Evaluation of DWP's Microsoft 365 Copilot Trial. Graded entry.
- Singla, A. et al. (2025). The state of AI. McKinsey. Graded entry.
- Vaccaro, M., Almaatouq, A. and Malone, T. (2024). When combinations of humans and AI are useful. Graded entry.
- Amodei, D. (2026). We Must Pace the Frontier. Graded entry.
- Microsoft AI (2026). Humanist AI in practice: a public consultation on our Code of Conduct for MAI models. Graded entry.
Related SuperSkills research#
On the difference between counting licences and measuring adoption, how do you measure AI adoption properly and usage theatre. On the record a decision has to leave, decision provenance. On whether a specialist director is the answer, should we appoint a director with AI expertise.
About this research#
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. He has run, grown, bought and advised businesses with AI in them. Findings are attributed to the studies and statements that produced them and kept separate from the interpretation. This is a living reference, reviewed and updated as significant new evidence appears.
Evidence review · SS-2026-244 · Graded against the published rubric
Hirji, R. (2026). How does a board know management's claims about AI are true?. The SuperSkills evidence base, SS-2026-244. https://thesuperskills.com/research/how-does-a-board-know-managements-claims-about-ai-are-true. Last reviewed 15 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work