← Research
Research

How does a board know management's claims about AI are true?

Almost every AI figure a board is shown is self-reported, by the people whose programme it measures. What independent assurance over AI would check, and why the disagreement rate is the number to ask for.

Last reviewed: 15 September 2026

Most AI figures presented to boards are self-reported by the people running the programme: time saved, adoption, productivity. Even government trials say so in their limitations chapters. What independent assurance over AI would check, the four documents a board can verify, and why the disagreement rate is the one number management cannot flatter. An evidence review by Rahim Hirji; every figure resolves to a graded entry in the evidence base that says what it does not show.

Questions this page answersAll 811 questions this research covers

Almost every AI figure a board is shown was produced by the people whose programme it measures. Time saved, licences activated, adoption, productivity: self-reported, by the enthusiasts, without a baseline. That is no accusation; the best-conducted trials say it about themselves. The UK Government Digital Service's cross-government Copilot experiment reported 26 minutes a day saved and stated that it could not identify how the saved time was spent; the Department for Work and Pensions' evaluation reported 19 minutes and listed, in its own limitations chapter, the absence of baseline data, self-selection towards enthusiasts and acquiescence bias on the time question. If the civil service says that about its own numbers, a board should assume it of a vendor deck.

The answer, in one line

Mostly it does not, because the claims are self-reported by the people whose programme they describe. The UK government's own Copilot trials reported 26 and 19 minutes a day saved and said in their limitations chapters that there was no control group, no baseline and self-selection of enthusiasts.

Share as a card

Why self-report is the norm#

Measuring what AI did to output quality or decision quality is expensive and slow, and measuring what people say it did is neither. McKinsey's 2025 survey of 1,491 organisations, the largest of its kind, measures self-reported EBIT impact and finds its 25 attributes explain a fifth of the variance. MIT NANDA's 95 per cent rests on 52 interviews. Gartner's 40 per cent is a forecast. Every number in circulation about enterprise AI is a claim by an interested party, and management's numbers are drawn from the same well. The board is not being lied to. It is being shown the only kind of figure that exists.

What assurance would check#

Independent assurance over AI, where AI takes part in consequential decisions, checks four documents rather than a dashboard. The allocation: which decisions the machine may inform, recommend or execute, which it may never own. The owners: a name against each. The capability floor: what the organisation keeps doing unaided so that somebody can still tell when the machine is wrong, and the evidence that they still can. And the disagreement rate: how often the humans reviewing machine output reach a different answer, measured rather than assumed. Vaccaro and colleagues' meta-analysis of 106 experiments found human and AI combinations underperforming the better of the two alone, with the losses in decision-making; a review that never disagrees is the mechanism by which that happens. A programme that cannot report the rate has no oversight, only a signature. One that reports zero has a rubber stamp.

The strongest argument for it came from the builders#

In September 2026 the chief executive of Anthropic asked, in public, for outside evaluators embedded inside AI companies with employee-level access, on the grounds that his own company's assurances about its own systems were not enough for anybody to rely on. Microsoft published rules for its models and opened them to six weeks of public consultation rather than announcing them as settled. Both are arguments, from the people with the most to lose by making them, that self-report is insufficient at the frontier. A board being told that its own management's self-report is sufficient is being asked to hold a lower standard than the laboratories now hold themselves to.

What nobody has measured#

There is no study of AI assurance and its effect on outcomes, and the four-document review described here is a proposal drawn from the model risk regime in banking rather than a tested practice. The evidence on this page is that the figures boards receive are self-reported, which their own producers confirm, and that the failures of AI programmes are consistently located in unmade decisions, which is what the four documents record. Whether reviewing them independently prevents failure is untested. That it makes the failure visible earlier is the claim.

What the board asks for#

Not a better dashboard. The four documents, once, and then whenever a system's scope changes; an annual external reading of them by somebody who does not report to the programme; and the disagreement rate every quarter, with the admission if it is not measured. The full list of questions is at what should a board ask about AI; what the audit trail has to contain is at how do you audit an AI-assisted decision; and the reason a governance pack does not answer this on its own is at AI governance versus AI leadership.

Key sources

On the difference between counting licences and measuring adoption, how do you measure AI adoption properly and usage theatre. On the record a decision has to leave, decision provenance. On whether a specialist director is the answer, should we appoint a director with AI expertise.

About this research#

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. He has run, grown, bought and advised businesses with AI in them. Findings are attributed to the studies and statements that produced them and kept separate from the interpretation. This is a living reference, reviewed and updated as significant new evidence appears.

How this research works  ·  Reviewed quarterly  ·  Found an error? Tell me and it is corrected on the page.

Evidence review · SS-2026-244 · Graded against the published rubric

Cite this page

Hirji, R. (2026). How does a board know management's claims about AI are true?. The SuperSkills evidence base, SS-2026-244. https://thesuperskills.com/research/how-does-a-board-know-managements-claims-about-ai-are-true. Last reviewed 15 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

How does a board know management's claims about AI are true?

Mostly it does not, because the claims are self-reported by the people whose programme they describe. The UK government's own Copilot trials reported 26 and 19 minutes a day saved and said in their limitations chapters that there was no control group, no baseline and self-selection of enthusiasts. A board verifies AI claims the way it verifies any other: by asking for the documents rather than the summary, and for numbers that cannot be flattered.

Does the board need independent assurance over AI?

Where AI takes part in consequential decisions, yes, and the shape is an annual external review of the allocation: the list of decisions the machine may make, the owners, the capability the organisation keeps unaided, and the disagreement rate. The chief executive of Anthropic asked in September 2026 for exactly that kind of embedded outside evaluation of his own company, which is the strongest available argument that self-report is not enough.

What is the one number a board should ask for?

The disagreement rate: how often a human reviewing machine output actually reaches a different answer. A programme that cannot report it has no oversight, only a signature. One that reports zero has a rubber stamp. It is the one figure that cannot be improved by enthusiasm.

In this hub

Definitions

The terms this field uses, defined against their primary sources.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

An annual external reading of the four documents by somebody who does not report to the programme. That reading is the engagement. Board advisory.

This argument is one a board usually meets for the first time in the room. There is the boards and leadership version, and the full range of topics and audiences.

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.