← Research
Research

Is it true that 95 per cent of AI pilots fail?

Three failure figures circulate in leadership rooms: 95, 80 and 40 per cent. Where each comes from, what each rests on, and the one thing all three agree about.

Last reviewed: 15 September 2026

The 95 per cent figure rests on 52 interviews in a preliminary MIT NANDA report; the 80 per cent is an estimate RAND cites rather than measures; the 40 per cent is a Gartner forecast with no published method. What each figure rests on, and the one finding all three share: the failures are in the organisation, not the model. An evidence review by Rahim Hirji; every figure resolves to a graded entry in the evidence base that says what it does not show.

Question this page answersAll 811 questions this research covers

Three numbers circulate in leadership rooms and each is quoted as though it were measured. Ninety-five per cent of AI pilots fail. Eighty per cent of AI projects fail. Forty per cent of agentic AI projects will be cancelled. The first rests on 52 interviews in a preliminary draft. The second is an estimate the report cites rather than one it made. The third is a forecast with no published method. None is a population figure and all three are quoted as one. What they do share, once the headline is set aside, is the finding that matters: in every source, the failures are in the organisation around the model and in the decisions nobody made before buying it, not in the model.

The answer, in one line

MIT NANDA's preliminary report The GenAI Divide, July 2025, which states that '95% of organizations are getting zero return' on generative AI.

Share as a card

Ninety-five per cent: 52 interviews and a draft#

MIT NANDA's report The GenAI Divide, July 2025, is the source, and its sentence is: 'Despite $30 to 40 billion in enterprise investment into GenAI, this report uncovers a surprising result in that 95% of organizations are getting zero return.' The method is 52 structured interviews, 153 survey responses gathered from senior leaders at four conferences, and a review of more than 300 publicly disclosed initiatives, over the first half of 2025. The report is marked as a preliminary version and is not peer-reviewed. Its own limitations line reads: the percentages 'reflect our interview sample of 52 organizations and may not represent broader market patterns'. Anyone quoting 95 per cent without the 52 is quoting a headline, and 'zero return' is the authors' phrase for no measurable P&L impact, which is not the same as no value.

What the report finds that almost nobody quotes is where the failures sit. The authors call it a learning gap: 'Most GenAI systems do not retain feedback, adapt to context, or improve over time', and users abandon tools that do not fit the workflow. Tools bought from specialist vendors reached deployment about twice as often as internal builds. And workers at over 90 per cent of the companies surveyed were using personal AI tools for work while only 40 per cent of companies had bought an official one, which means the adoption the leadership had not decided on was running ahead of the adoption it had.

Eighty per cent: an estimate, cited#

RAND's August 2024 report on why AI projects fail opens with 'by some estimates, more than 80 percent of AI projects fail', twice the rate of ordinary IT projects. RAND did not measure that; it cites it and moves on to its own contribution, which is 65 semi-structured interviews with practitioners and five root causes. The first listed is leadership-driven failure: misunderstanding what problem the project is for, or optimising the wrong metric. Then data, then bottom-up failures where the technology was chosen before the problem, then infrastructure, then immature technology. Four of the five are decisions; the fifth is the one everybody blames.

Forty per cent: a forecast#

Gartner's press release of 25 June 2025 predicts that 'over 40% of agentic AI projects will be canceled by the end of 2027, due to escalating costs, unclear business value or inadequate risk controls'. No method is published for the forecast. The supporting poll is of 3,412 webinar attendees, who are people already interested enough to attend, and Gartner sells advice on the projects it is forecasting. The useful part is the three reasons, because each is a decision a leadership team could have made before it bought anything: what it would cost, what it was for, and who controlled it. Gartner also estimates that of the thousands of vendors claiming agentic products, about 130 are real; the rest it calls agent washing.

What the three agree on#

Set the percentages aside, since none survives contact with its own footnote, and the three sources say one thing from three directions. MIT locates failure in integration, learning and workflow. RAND locates it in unclear objectives and wrong metrics. Gartner locates it in unpriced cost, undefined value and absent controls. None locates it in the model. McKinsey's 2025 survey, from the other end, found the attributes most associated with profit from generative AI to be the chief executive owning the rules and the redesign of workflows, held by 28 and 21 per cent of organisations. The failures and the successes are described in the same vocabulary, the vocabulary of decisions.

What nobody has measured#

There is no census of enterprise AI projects with a defined failure criterion and a population from which the rate could be estimated. Every figure here is either a small interview sample, a cited estimate or a forecast, and every one comes from an organisation with something to sell to the people who quote it. The UK government's own Copilot trials, which did attempt measurement, produced self-reported time savings of 19 and 26 minutes a day with no control group and no measure of output quality, and are graded accordingly. The true failure rate is unknown. The direction of the causes is not.

The question to ask instead#

A leadership team that has heard the 95 per cent figure should not ask whether it is true. It should ask which of the three named causes its own programme has already settled in writing: what the project is for and how that will be measured, how it fits the work people actually do, and who has the authority to stop it. A programme with those three written down is not in the 95 per cent, whatever the 95 per cent turns out to be. The rules for writing them are at rules before tools, and the case that this is leadership work rather than technology work is at AI leadership.

Key sources

On measuring adoption rather than counting licences, how do you measure AI adoption properly. On usage that looks like adoption and is not, usage theatre. On what goes wrong in transformations, common AI transformation challenges. On who should hold the decisions the three sources name, AI leadership.

About this research#

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. He has run, grown, bought and advised businesses with AI in them. Findings are attributed to the studies and statements that produced them and kept separate from the interpretation. This is a living reference, reviewed and updated as significant new evidence appears.

How this research works  ·  Reviewed quarterly  ·  Found an error? Tell me and it is corrected on the page.

Evidence review · SS-2026-242 · Graded against the published rubric

Cite this page

Hirji, R. (2026). Is it true that 95 per cent of AI pilots fail?. The SuperSkills evidence base, SS-2026-242. https://thesuperskills.com/research/is-it-true-that-95-per-cent-of-ai-pilots-fail. Last reviewed 15 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

Where does the 95 per cent figure come from?

MIT NANDA's preliminary report The GenAI Divide, July 2025, which states that '95% of organizations are getting zero return' on generative AI. The figure rests on 52 structured interviews, 153 survey responses gathered at conferences and a review of 300 public initiatives, and the authors write that the percentages 'reflect our interview sample of 52 organizations and may not represent broader market patterns'.

Where does the 80 per cent figure come from?

RAND's August 2024 report on the root causes of AI project failure, which states that 'by some estimates, more than 80 percent of AI projects fail', twice the rate of ordinary IT projects. RAND cites the estimate; it did not measure it. The report's own contribution is 65 interviews and five root causes, of which leadership-driven failure is listed first.

Where does the 40 per cent figure come from?

Gartner's press release of 25 June 2025 forecasting that over 40 per cent of agentic AI projects will be cancelled by the end of 2027 through escalating costs, unclear business value or inadequate risk controls. It is a forecast with no published method, supported by a poll of 3,412 webinar attendees.

So do most AI pilots fail?

Probably, and nobody has measured it well. What the three sources agree on is where the failures sit: in integration, workflow, unclear objectives and absent controls, which are organisational and leadership decisions, not model quality.

In this hub

Definitions

The terms this field uses, defined against their primary sources.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Whatever the true rate, the three causes every source names are decisions a leadership team can settle before it buys. Settling them for one programme is the engagement. Board advisory.

This argument is one a board usually meets for the first time in the room. There is the boards and leadership version, and the full range of topics and audiences.

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.