Three numbers circulate in leadership rooms and each is quoted as though it were measured. Ninety-five per cent of AI pilots fail. Eighty per cent of AI projects fail. Forty per cent of agentic AI projects will be cancelled. The first rests on 52 interviews in a preliminary draft. The second is an estimate the report cites rather than one it made. The third is a forecast with no published method. None is a population figure and all three are quoted as one. What they do share, once the headline is set aside, is the finding that matters: in every source, the failures are in the organisation around the model and in the decisions nobody made before buying it, not in the model.
The answer, in one line
MIT NANDA's preliminary report The GenAI Divide, July 2025, which states that '95% of organizations are getting zero return' on generative AI.
Ninety-five per cent: 52 interviews and a draft#
MIT NANDA's report The GenAI Divide, July 2025, is the source, and its sentence is: 'Despite $30 to 40 billion in enterprise investment into GenAI, this report uncovers a surprising result in that 95% of organizations are getting zero return.' The method is 52 structured interviews, 153 survey responses gathered from senior leaders at four conferences, and a review of more than 300 publicly disclosed initiatives, over the first half of 2025. The report is marked as a preliminary version and is not peer-reviewed. Its own limitations line reads: the percentages 'reflect our interview sample of 52 organizations and may not represent broader market patterns'. Anyone quoting 95 per cent without the 52 is quoting a headline, and 'zero return' is the authors' phrase for no measurable P&L impact, which is not the same as no value.
What the report finds that almost nobody quotes is where the failures sit. The authors call it a learning gap: 'Most GenAI systems do not retain feedback, adapt to context, or improve over time', and users abandon tools that do not fit the workflow. Tools bought from specialist vendors reached deployment about twice as often as internal builds. And workers at over 90 per cent of the companies surveyed were using personal AI tools for work while only 40 per cent of companies had bought an official one, which means the adoption the leadership had not decided on was running ahead of the adoption it had.
Eighty per cent: an estimate, cited#
RAND's August 2024 report on why AI projects fail opens with 'by some estimates, more than 80 percent of AI projects fail', twice the rate of ordinary IT projects. RAND did not measure that; it cites it and moves on to its own contribution, which is 65 semi-structured interviews with practitioners and five root causes. The first listed is leadership-driven failure: misunderstanding what problem the project is for, or optimising the wrong metric. Then data, then bottom-up failures where the technology was chosen before the problem, then infrastructure, then immature technology. Four of the five are decisions; the fifth is the one everybody blames.
Forty per cent: a forecast#
Gartner's press release of 25 June 2025 predicts that 'over 40% of agentic AI projects will be canceled by the end of 2027, due to escalating costs, unclear business value or inadequate risk controls'. No method is published for the forecast. The supporting poll is of 3,412 webinar attendees, who are people already interested enough to attend, and Gartner sells advice on the projects it is forecasting. The useful part is the three reasons, because each is a decision a leadership team could have made before it bought anything: what it would cost, what it was for, and who controlled it. Gartner also estimates that of the thousands of vendors claiming agentic products, about 130 are real; the rest it calls agent washing.
What the three agree on#
Set the percentages aside, since none survives contact with its own footnote, and the three sources say one thing from three directions. MIT locates failure in integration, learning and workflow. RAND locates it in unclear objectives and wrong metrics. Gartner locates it in unpriced cost, undefined value and absent controls. None locates it in the model. McKinsey's 2025 survey, from the other end, found the attributes most associated with profit from generative AI to be the chief executive owning the rules and the redesign of workflows, held by 28 and 21 per cent of organisations. The failures and the successes are described in the same vocabulary, the vocabulary of decisions.
What nobody has measured#
There is no census of enterprise AI projects with a defined failure criterion and a population from which the rate could be estimated. Every figure here is either a small interview sample, a cited estimate or a forecast, and every one comes from an organisation with something to sell to the people who quote it. The UK government's own Copilot trials, which did attempt measurement, produced self-reported time savings of 19 and 26 minutes a day with no control group and no measure of output quality, and are graded accordingly. The true failure rate is unknown. The direction of the causes is not.
The question to ask instead#
A leadership team that has heard the 95 per cent figure should not ask whether it is true. It should ask which of the three named causes its own programme has already settled in writing: what the project is for and how that will be measured, how it fits the work people actually do, and who has the authority to stop it. A programme with those three written down is not in the 95 per cent, whatever the 95 per cent turns out to be. The rules for writing them are at rules before tools, and the case that this is leadership work rather than technology work is at AI leadership.
Key sources
- Challapally, A., Pease, C., Raskar, R. and Chari, P. (2025). The GenAI Divide: State of AI in Business 2025. MIT NANDA, preliminary report. Graded entry.
- Ryseff, J., De Bruhl, B. F. and Newberry, S. J. (2024). The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed. RAND. Graded entry.
- Gartner (2025). Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027. Graded entry.
- Singla, A. et al. (2025). The state of AI. McKinsey. Graded entry.
- Government Digital Service (2025). Microsoft 365 Copilot Experiment: Cross-Government Findings Report. Graded entry.
- Department for Work and Pensions (2026). An Evaluation of DWP's Microsoft 365 Copilot Trial. Graded entry.
Related SuperSkills research#
On measuring adoption rather than counting licences, how do you measure AI adoption properly. On usage that looks like adoption and is not, usage theatre. On what goes wrong in transformations, common AI transformation challenges. On who should hold the decisions the three sources name, AI leadership.
About this research#
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. He has run, grown, bought and advised businesses with AI in them. Findings are attributed to the studies and statements that produced them and kept separate from the interpretation. This is a living reference, reviewed and updated as significant new evidence appears.
Evidence review · SS-2026-242 · Graded against the published rubric
Hirji, R. (2026). Is it true that 95 per cent of AI pilots fail?. The SuperSkills evidence base, SS-2026-242. https://thesuperskills.com/research/is-it-true-that-95-per-cent-of-ai-pilots-fail. Last reviewed 15 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work