Governance is the machinery for checking decisions that have been made. Leadership is making them. The distinction sounds like a definition and behaves like a diagnosis, because most organisations have spent two years building AI governance, meaning policies, risk tiers, review councils, model inventories and audit trails, and are now waiting for it to produce the decisions it was built to check. It cannot. A governance framework can only test a deployment against rules somebody has set. If nobody has decided which decisions the machine may make, where a human must remain and what the organisation's people must stay capable of, the framework has nothing to test against and passes everything. That is how an organisation ends up compliant and drifting at the same time.
The answer, in one line
Governance is the machinery for checking, recording and controlling AI decisions once they have been made: policies, risk tiers, review councils, audit trails, model inventories.
What each one is#
AI governance is the set of structures that review, record and control AI use once it exists. A policy people can find. Risk tiers, with more human review as the stakes rise. A council that can approve, pause or stop. A register of live systems with a purpose, an owner and a review date. A log that would let somebody reconstruct why a decision was made. Article 12 of the EU AI Act requires the log for high-risk systems; Article 14 requires oversight that a human can actually exercise. Governance is necessary, it is well understood, and firms exist to build it.
AI leadership is the allocation of judgement. Deciding, in advance and in writing, which decisions a machine may inform, which it may recommend, which it may execute and which stay human; who is accountable for each; and what the organisation's people must remain capable of doing so that the humans in charge can still exercise the judgement the role needs. It is the input governance was built to check. The definition and the evidence are at AI leadership.
Why the first does not produce the second#
A review council reviews what is put in front of it. A risk tier classifies a system somebody has already chosen to deploy. An audit trail records a decision after it is made. Each is downstream of a choice about scope, and none of them makes that choice. In practice the choice gets made by whoever configures the system, which is usually a vendor or an implementation team optimising for adoption, and the governance machinery then certifies the result. The organisation has a compliant record of decisions it never took.
The evidence on failure reads the same way. RAND's five root causes of AI project failure list leadership-driven failure first: misunderstanding what the project was for, or optimising the wrong metric. Gartner's three reasons for the cancellations it forecasts are unpriced cost, undefined value and inadequate risk controls; only the third is a governance failure, and that one governance frameworks are best at catching. MIT NANDA locates failure in workflow and integration, which no council reviews. McKinsey's 2025 survey found the chief executive's personal oversight of AI governance to be among the attributes most associated with reported profit, which is the two joined: a leader making the decisions, and the machinery checking them.
The test that tells them apart#
Ask the organisation for two documents. The first is the governance pack: policy, tiers, council terms of reference, system register. Most organisations of any size can produce it. The second is the allocation: the list of consequential decisions with each marked inform, recommend, execute or never, a named owner against each, the work the organisation keeps unaided, and the rate at which human reviewers disagree with machine output. Very few can produce it, and the ones that cannot have governance without leadership. The council has been meeting; the decisions it exists to check have been made by default.
The disagreement rate is the sharpest single test, because it belongs to both. Governance can require that it be measured. Only leadership can decide what rate is acceptable and what happens when it falls to zero, the point at which a human in the loop has become a signature on the loop. Vaccaro and colleagues' meta-analysis of 106 experiments found human and AI combinations underperforming the better of the two alone, with the losses in decision-making; a review that never disagrees is the mechanism.
What nobody has measured#
No study separates organisations with strong governance and weak leadership from the reverse and follows their outcomes, and the two are rarely measured apart. The argument here is from the logic of the structures: what each can and cannot do by construction. The supporting evidence is on where failures come from, which is consistently the unmade decision rather than the unreviewed one.
What a board asks for#
From governance, the register, the tiers, the audit trail and the incident process, quarterly. From leadership, the allocation, the owners, the capability floor and the disagreement rate, once, in writing, and then whenever a system's scope changes. A board that is shown the first and not the second is being shown the machinery and not the decisions, and should say so. The questions a board asks are at what should a board ask about AI; what the oversight looks like when it works is at what board oversight of AI looks like; the ten rules for writing the allocation are at rules before tools.
Key sources
- Ryseff, J., De Bruhl, B. F. and Newberry, S. J. (2024). The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed. RAND. Graded entry.
- Gartner (2025). Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027. Graded entry.
- Challapally, A. et al. (2025). The GenAI Divide: State of AI in Business 2025. MIT NANDA. Graded entry.
- Singla, A. et al. (2025). The state of AI. McKinsey. Graded entry.
- Vaccaro, M., Almaatouq, A. and Malone, T. (2024). When combinations of humans and AI are useful. Nature Human Behaviour. Graded entry.
- Regulation (EU) 2024/1689, Annex III. Graded entry.
Related SuperSkills research#
The definition is at AI leadership. The governance side in detail: meaningful human oversight, how do you audit an AI-assisted decision and decision provenance. The leadership side: how AI decision rights should be allocated and drift versus design.
About this research#
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. He has run, grown, bought and advised businesses with AI in them. Findings are attributed to the studies and statements that produced them and kept separate from the interpretation. This is a living reference, reviewed and updated as significant new evidence appears.
Evidence review · SS-2026-239 · Graded against the published rubric
Hirji, R. (2026). AI governance versus AI leadership. The SuperSkills evidence base, SS-2026-239. https://thesuperskills.com/research/ai-governance-versus-ai-leadership. Last reviewed 15 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work