AI leadership is the allocation of judgement. Every organisation now has an AI strategy, in the sense of a plan for what to buy, build and deploy. Far fewer have anyone deciding which decisions the machines may make, which stay with people, who answers for each, and what the people must remain capable of so that the humans nominally in charge still are. That second set of decisions does not arrive on an agenda by itself, and nothing fails while it goes unmade. It is nonetheless the whole of what leading an organisation through AI consists of. The tools are the easy half. The rules are the job.
The answer, in one line
AI leadership is the allocation of judgement: deciding, in advance and in writing, which decisions a machine may inform, which it may recommend, which it may execute, and which stay human; who is accountable for each; and what the organisation's people must remain capable of doing so that the humans in charge can still exercise the judgement the role needs.
Definition#
AI leadership: the allocation of judgement. Deciding, in advance and in writing, which decisions a machine may inform, which it may recommend, which it may execute, and which stay human; who is accountable for each; and what the organisation's people must remain capable of doing so that the humans in charge can still exercise the judgement the role needs. It does not require technical fluency. It requires the capacity to judge what the technical people build, and the standing to say no to it. Rahim Hirji's reading of a phrase in common use, set out in SuperSkills (Kogan Page, 2026) and in Rules Before Tools (August 2025). No claim of first use is made for the phrase; the definition is the claim.
The largest recurring survey of corporate AI use makes the point in its own numbers. Of the twenty-five things McKinsey tested against reported profit from generative AI in 2025, the one most associated with money was the chief executive personally overseeing the rules for how AI is developed and deployed, and the one with the biggest effect was the redesign of workflows. Twenty-eight per cent of organisations had the first. Twenty-one per cent had done the second. Both are leadership acts, neither is a technology act, and most organisations have neither.
Three kinds of work, and one of them is unowned#
Watch any organisation bring in AI and three kinds of work appear. Transformation work: the process redesign, the change programme, the operating model, owned by a transformation lead or a chief operating officer. Technical work: the data, the models, the integration, the security, owned by a chief technology officer, a chief information officer or the vendors. Both are real disciplines with real owners, and both are well served.
The third is the judgement around both. What the organisation should allow these systems to do. What its people must stay capable of doing once the systems do most of it. Who answers when a machine took part in a decision that went wrong. In most organisations nobody owns this. It is treated as a property that will emerge from the first two if they are done well enough, and it does not, because neither the transformation lead nor the technologist is paid to decide what the organisation should refuse to automate. That third box is leadership. It is also, by definition, the part that cannot be bought from the firms that sell the other two.
What the failures have in common#
The failure figures that circulate in leadership rooms are mostly worse than they look and all point the same way. MIT NANDA's 2025 report is quoted everywhere as showing that 95 per cent of enterprise AI pilots return nothing; that figure rests on 52 interviews and the authors say so. What the same report finds, and almost nobody quotes, is where the failures sit: not in the models but in the organisation around them, in integration, learning and workflow, while workers at nine in ten of the companies surveyed were using AI on their own initiative and only four in ten companies had bought it for them. Gartner's forecast that over 40 per cent of agentic AI projects will be cancelled by the end of 2027 has no published method, and its three stated reasons are the useful part: escalating costs, unclear business value and inadequate risk controls. Each is a decision a leadership team could have made before it bought anything, and did not.
The two most-cited corporate cases make the same point from opposite directions. In 2024 Klarna announced that its AI assistant was doing the work of 700 customer service agents. In May 2025 its chief executive, Sebastian Siemiatkowski, told Bloomberg that cost had been 'a too predominant evaluation factor' and that 'what you end up having is lower quality', and the company began recruiting people again. The technology had done what it was asked to. The decision about what to ask it had been made on one variable. At IBM in the same week, Arvind Krishna described replacing a few hundred human resources staff with AI agents and then choosing to put the money into engineers and salespeople, so that 'our total employment has actually gone up'. Both are chief executives describing their own companies and neither published the figures. Read together, they show that the outcome of the same automation depends on a choice made above it, and that the choice is the leader's.
You do not need to be technical. You need to be able to judge.#
The most common reason senior people give for delegating AI to the technologists is that they do not understand the technology. The decisions in the third box are not technical decisions. Whether a credit decision may be executed by a model or only recommended by one is a question about accountability and appetite. Whether the organisation keeps some underwriting unaided so that its underwriters can still check the model is a question about capability. Who has the authority to stop an agent, and how often a human reviewing machine output actually reaches a different answer, are questions about governance. A leader who could not build any of these systems can decide all of these things, and a leader who could build them has no particular advantage in deciding them well.
Fluency helps and should be worked at. What cannot be delegated is the judging. Banking's model risk regime, the most mature oversight practice any industry has for machine-produced numbers, rests on effective challenge by people with three things: independence, standing and expertise. Independence is structural and can be arranged. Standing is political and can be granted. Expertise is built by doing the work, and an organisation that has automated the work its challengers learned on has removed the third leg while keeping the other two. Lisanne Bainbridge wrote that down about process control in 1983. It is the mechanism underneath capability debt, and it means the leadership decision covers what to keep people doing, so that somebody can still tell when the machine is wrong, as much as what to allow the machine to do today.
Rules before tools#
In August 2025, before the phrase became fashionable, the argument for putting the rules before the roadmap was set out in ten of them: start with tasks; redesign the system before automating it; put a human in command of every initiative; make decisions legible where they touch people or money; measure impact and harm together; strengthen the data backbone; build guardrails in by default, because clear guardrails speed you up; treat capacity as a strategy; protect dignity by design; and educate for leverage, the tenth rule's own title. The line was 'the next model won't save you. A better system will.' A year on, that reads as the difference between the organisations in McKinsey's 21 per cent and everybody else.
In practice the rules reduce to one document that most leadership teams have never written. A list of the decisions the organisation makes, with each marked as one the machine may inform, recommend or execute, or one it may never own. An owner against each. The capabilities the organisation intends to keep practising, unaided, and where. And a number: how often the humans reviewing machine output disagree with it, because a review that never disagrees is not a review. Writing that list is usually the most useful hour a leadership team spends on AI, and it costs nothing but the argument.
The same argument, at the scale of the industry#
In September 2026 the people building the frontier models said the same thing about themselves. Dario Amodei asked for outside evaluators embedded inside the AI companies with employee-level access. Microsoft published rules under which its models 'will never resist human interruption, correction, or shutdown'. Both remedies are people inside the loop with real access and the capability to use it, and neither document says where those people come from or what happens when the humans nominally in charge have stopped practising the judgement the role needs. That is the third box again, at civilisational scale. The dated record of that week is on the Third Way page. An organisation cannot settle the frontier question. It can settle its own.
What nobody has measured#
There is no study that measures the allocation of judgement inside organisations and relates it to outcomes, so the argument on this page is assembled from adjacent evidence and should be read that way. McKinsey's correlation explains a fifth of the variance in a self-reported outcome, and chief executives who already make money from AI may simply be the ones who take an interest in its rules. The MIT figure is a small interview sample in a preliminary draft. Gartner is a forecast. Klarna and IBM are two men describing their own companies. The estate's own instrument, the Drift versus Design Matrix, is a structured self-assessment and not an independent measure. What can be said is that every source, from every motive, locates the failure in decisions rather than in models, and that none of them finds an organisation that succeeded by leaving those decisions unmade.
What a leadership team decides this quarter#
Five things, each with a decision at the end of it. Which decisions AI may recommend, which it may execute, and which require a human to make them, written down. Who owns each, by name. Where the organisation keeps some work unaided so that its people can still check the machine, and how it will know if they no longer can. How often its reviewers disagree with machine output, measured. And what the executive team will understand personally rather than delegate to a steering committee, because a board that cannot interrogate its own AI decisions has delegated more than it intended. The ongoing version of this work is at AI adviser to CEOs, boards and leadership teams; the room version is the AI leadership keynote.
Key sources
- Singla, A., Sukharevsky, A., Yee, L., Chui, M. and Hall, B. (2025). The state of AI: How organizations are rewiring to capture value. McKinsey and Company, 12 March 2025. Graded entry.
- Challapally, A., Pease, C., Raskar, R. and Chari, P. (2025). The GenAI Divide: State of AI in Business 2025. MIT NANDA, preliminary report. Graded entry.
- Gartner (2025). Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027. Press release, 25 June 2025. Graded entry.
- Siemiatkowski, S. (2025), reported in Klarna changes its AI tune and again recruits humans for customer service. CX Dive, 9 May 2025, from a Bloomberg interview. Graded entry.
- Krishna, A. (2025), reported in IBM CEO: layoffs due to AI led to 'more investment' in other roles. PYMNTS, 6 May 2025, from a Wall Street Journal interview. Graded entry.
- Vaccaro, M., Almaatouq, A. and Malone, T. (2024). When combinations of humans and AI are useful. Nature Human Behaviour. Graded entry.
- Boston Consulting Group (2026). When Everyone Uses AI, Companies Risk Losing Critical Skills. Graded entry.
- Bainbridge, L. (1983). Ironies of automation. Automatica. Graded entry.
- Microsoft AI (2026). Humanist AI in practice: a public consultation on our Code of Conduct for MAI models. Graded entry.
- Amodei, D. (2026). We Must Pace the Frontier. Graded entry.
- Hirji, R. (2025). Rules Before Tools. Box of Amazing, 17 August 2025.
Related SuperSkills research#
AI leadership is the decision layer above ideas developed elsewhere on this site: drift versus design, which is the difference between allocating judgement and letting it move; capability debt, the bill for leaving the capability question unmade; meaningful human oversight, on what a human in the loop has to be able to do before it counts as governance; AI and human judgement; and synthetic seniority, the question a chief executive tends to recognise fastest. For the dated record of the September 2026 frontier debate, see neither hype nor doom.
About this research#
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. He has run, grown, bought and advised businesses with AI in them, and this page argues from that seat as well as from the studies. Findings are attributed to the studies that produced them and kept separate from the interpretation. The definition of AI leadership on this page is his; the phrase is in common use and no claim of first use is made. This is a living reference, reviewed and updated as significant new evidence appears.
Evidence review · SS-2026-238 · Graded against the published rubric
Hirji, R. (2026). AI leadership. The SuperSkills evidence base, SS-2026-238. https://thesuperskills.com/research/ai-leadership. Last reviewed 15 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work