← Research
Research

Which AI investments should we stop?

The obstacle is not analysis. Four decades of research says the obstacle is that the people who started a programme are the ones asked whether to end it.

Last reviewed: 4 September 2026

The escalation literature, its four-phase model of how organisations climb back down, why the most quoted AI failure statistics do not survive being traced to source, and a stopping test that does not require a counterfactual nobody can build.

Question this page answersQuestion this page partly answersAll 811 questions this research covers

Stop the programmes where three things are true at once: nobody can now state what the work was supposed to change, the only evidence of benefit is self-reported or supplied by the seller, and the person who would have to recommend stopping is the person who started it. Those conditions describe escalation, which has fifty years of research behind it, rather than a difficult project, which has none. What the test deliberately does not require is a counterfactual estimate of the programme's value. Almost no organisation can construct one, and the wait for it has become the most respectable way of not deciding.

The answer, in one line

Stop the ones where nobody can now state what the programme was supposed to change, where the only evidence of benefit is self-reported or supplied by the vendor, and where the sponsor and the assessor are the same person.

Share as a card

The founding case study of runaway IT spending was an expert system#

Mark Keil's 1995 study in MIS Quarterly is the paper that put project escalation into information systems research. It is a single longitudinal case, built from 111 interviews, 19 observed meetings and more than 350 collected documents inside a large computer manufacturer he calls CompuSys. The project he follows, CONFIG, was an expert system designed to help sales representatives produce error-free configurations before quoting a price. In Keil's words, "after more than a decade of development and tens of millions of dollars, the CONFIG project was eventually terminated at the end of 1992".

The business case history is the instructive part. A 1982 analysis put the net present value at $43.9 million against projected operating and development costs of $10.4 million. A 1985 re-forecast raised the figure to $55.7 million. A 1987 analysis put it at at least $41.1 million. By 1991 the group's annual operating budget was around $45 million. Each re-forecast was produced by people who believed in the thing and had been right about it before, and each one arrived at a number large enough to justify continuing.

So the canonical study of a technology programme that could not be stopped is a study of an artificial intelligence programme. That is worth sitting with. The literature this page draws on is not being applied to AI by analogy. It began there, thirty years ago, and then the field of AI investment forgot it.

A measured prevalence, and a mechanism that has been reproduced since 1976#

Keil went on to measure how common the pattern is. Keil, Mann and Rai surveyed information systems audit and control professionals, designing the instrument to capture projects that did not escalate as well as those that did, and report that "the results of our research suggest that between 30% and 40% of all IS projects exhibit some degree of escalation". Two caveats belong with that number. It is a retrospective survey of auditors rather than a random sample of projects, and "some degree of escalation" is a soft threshold. The published sample size sits inside the full text, which could not be opened for this page; the prevalence figure is quoted from the abstract, and that is stated here rather than glossed over.

The mechanism underneath it is older and cleaner. Barry Staw ran 240 business students through a role-played corporate funding decision in a two-by-two design crossing personal responsibility against decision consequences. Participants who had personally made the earlier investment allocated an average of $11.08 million to the division they had chosen, against $8.89 million where the earlier choice had been made by another officer. Where their own choice had subsequently declined, the figure rose to $13.07 million. Negative consequences drew more money than positive ones, at $11.20 million against $8.77 million. Both main effects were significant and so was the interaction.

Staw's study is a paper exercise with undergraduates and no real money, and it should not be asked to carry more than it can. What it establishes is narrow and durable: responsibility for the original decision changes the subsequent one, in a known direction, before any question of competence arises. The general effect it belongs to, the sunk cost effect, was named by Arkes and Blumer in 1985. Their specific experimental figures are widely quoted and are not quoted here, because the full text could not be read at source for this page and the secondary versions in circulation disagree with one another.

Denver Airport gave us the only map of the climb back down#

Escalation research is large. De-escalation research is not, and Montealegre and Keil said so at the time: "while prior research has shown that managers can easily become locked into a cycle of escalating commitment to a failing course of action, there has been comparatively little research on de-escalation, or the process of breaking such a cycle". Their study of the baggage handling system at Denver International Airport produced the model that is still the field's reference. They describe de-escalation as a four-phase process: "(1) problem recognition, (2) re-examination of prior course of action, (3) search for alternative course of action, and (4) implementing an exit strategy".

The four phases are useful for a reason that is easy to miss. They separate noticing from being permitted to act. Organisations reach phase one constantly. Somebody in the room always knows. What they mostly do not reach is phase two, because re-examining the prior course of action means asking the person who chose it to say it was wrong, in front of the people who approved the funding. That is a governance problem wearing the costume of an analytical one, and no amount of better measurement touches it.

It also explains why "let us gather six more months of data" is such a popular answer. It is phase-one activity: defensible, costing nobody any standing, and impossible to tell apart from phase two until a year has gone. Montealegre and Keil's model is inductive, built from a single case, and it has never been tested for how often de-escalation succeeds. Treat it as a description of the route, not evidence that the route is usually taken.

The failure rate everybody quotes traces to a magazine article#

A page telling organisations to stop things ought to be able to say how often AI programmes fail. It cannot, and the reasons are worth publishing, because the numbers in circulation are doing real damage to real capital decisions.

RAND's 2024 report on the root causes of failure in artificial intelligence projects, built from interviews with 65 experienced data scientists and engineers, opens by observing that "by some estimates, more than 80 percent of AI projects fail. This is twice the already-high rate of failure in corporate information technology (IT) projects that do not involve AI." Follow the endnote. It points to a press article. Follow the second endnote, attached to the comparison, and it points to a business magazine piece. The most cited statistic about AI project failure, traced through the most careful organisation that repeats it, resolves to journalism. RAND hedge it correctly as "by some estimates"; everybody who quotes RAND drops the hedge.

The other famous figure, that 95 per cent of organisations get zero return from generative AI, comes from a document titled The GenAI Divide: State of AI in Business 2025, attributed to MIT NANDA and dated July 2025. Its own cover page describes it as "Preliminary Findings", the file is versioned 0.1, it carries a disclaimer stating that the views "do not reflect the positions of any affiliated employers", and no MIT domain hosts it. Its method is a review of publicly disclosed initiatives, 52 structured interviews and 153 survey responses "collected across four major industry conferences". The 95 per cent refers to organisations reporting no measured profit-and-loss return, which is not the same claim as the one that travels. This estate has refused that figure before and refuses it again.

Gartner and S&P Global figures on AI abandonment circulate constantly. Neither is readable at source without payment and neither publishes a method, so neither can be checked. That is a statement about verifiability rather than about quality, and it explains their absence here.

The one national statistic, and what it is actually counting#

The US Census Bureau's Business Trends and Outlook Survey is the best firm-level instrument on AI use anywhere: roughly 200,000 businesses sampled bi-weekly, nationally representative, weighted, method fully published. It does not ask whether a firm stopped using AI. It asks about current use and expected use in six months, so discontinuation can only be inferred from the gap between them.

Bonney and colleagues report that "a large fraction of the firms (67.9%) that currently use AI also expect to use in the future. However, a non-trivial fraction (14.5%) of the current users do not expect to use in the future, and another 17.6% don't know whether they will. Thus, about one in seven (and possibly more) of the current AI users may 'de-adopt' in the future." Their explanation is the interesting one: "de-adoption may occur if such experimentation does not yield anticipated benefits or organizational synergies."

Read it precisely. One in seven is a stated intention, not an observed exit, and the survey covers firm-level use rather than programmes or investments, so it cannot tell you that anybody cancelled anything. What it does establish is that stepping back from AI is a normal outcome at scale in the American economy, reported by firms themselves, at a national statistical agency, in a period when saying so publicly was unfashionable. Anyone told that stopping would be an admission of failure should be shown that sentence.

Three questions that do not require the counterfactual#

The standard demand made of anyone proposing to stop an AI programme is that they prove it is not working. That demand cannot be met. Establishing what a programme was worth requires a counterfactual, and the estate has already set out how long the signal takes to arrive and why the fastest available measures are the least reliable. Requiring proof of failure before permitting a stop is therefore not a high standard. It is an unmeetable one, which is what makes it so useful to whoever wants to continue.

These three questions are answerable this week, from documents that already exist. They are built on Montealegre and Keil's phase two, the re-examination that organisations skip, and their purpose is to make that phase possible without anyone having to concede that they were wrong.

A fourth question follows from this research rather than from the escalation literature. Nobody asks it. What can your people no longer do unaided? A programme can be worth stopping and still have changed the organisation permanently, because the practice it displaced does not come back when the licence lapses. That is capability debt. A cancellation therefore does not put you back where you started. Stopping is cheap. Rebuilding the reps is not. The estate's argument about missed reps applies with more force here than anywhere, because the people who stopped practising during the pilot are the people who will be asked to run the manual process afterwards.

Where this argument came from#

This is not a new position for this research. In Time-as-a-Service in 2023, Rahim Hirji argued that the emerging business model was not software but time, and that firms would monetise the time they gave back. Hours saved became the standard enterprise AI metric roughly two years later, which is a prediction worth recording precisely because it makes the present argument uncomfortable: the metric he saw coming is the one this page says will not answer a capital question.

In AI for AI's Sake in 2024 he argued that bolting AI onto products that do not need it degrades them, and is usually marketing or patent defence rather than value. In Rules Before Tools, published on 17 August 2025, the position hardened into a working method. "The part most people don't want to hear? The boring work is the valuable work." Among the ten rules is one about measurement that belongs on this page in full: pilot with a small share of users or one region, "track one business metric and one safety metric, and know when to roll back." Designing the exit at the start is the single cheapest form of de-escalation available, because at that point nobody has anything invested in the answer.

And in The Decision You Never Made in 2025 he set out the condition that makes stopping so hard: the consequential choices about AI in most organisations were never made by anyone in particular. They accumulated out of individual convenience. A programme nobody decided to start is a programme nobody has standing to end, which is drift rather than design arriving in the capital budget.

The number this page refuses to give you#

There is no defensible base rate for AI programme failure and this page does not supply one. Keil, Mann and Rai's 30 to 40 per cent covers information systems projects in 2000, not AI programmes in 2026. The Census 14.5 per cent is an intention about firm-level use, not a programme outcome. RAND's 65 interviews are a root-cause study, and RAND do not claim otherwise. Anybody quoting a single percentage for AI project failure is quoting something that has not been measured.

Nor is there peer-reviewed work applying escalation theory to AI or machine learning investment. That was searched for across 2020 to 2026 and none was found. The nearest item is a preprint on whether language models themselves display the sunk cost bias, which is a different question with the machine as the subject. So the framework on this page is transferred from information systems research, and the transfer is an argument rather than a finding.

One more limit, and it cuts against the page's own recommendation. Escalation is not the only reason programmes continue, and the research does not show that escalated projects would have been better off terminated. It shows their outcomes were worse. Some long, expensive, unpopular programmes are correct. The three questions above are designed to make a decision possible, not to make it come out one way.

Key sources

On the measurement problem underneath the capital question, how long before you know if an AI investment worked, how to measure AI adoption properly and usage theatre. On stopping for reasons other than money, deployment is not a ratchet and how to design a stop button people will use. On what a cancellation leaves behind, capability debt, missed reps and preserving capability across vendors. On who should be holding the question, what a board should ask about AI and who should own AI strategy.

About this research#

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. Escalation of commitment, the sunk cost effect and de-escalation are established concepts in organisational behaviour and information systems research, credited above to the people who developed them. Nothing on this page is a SuperSkills coinage. Missed reps is his. Capability debt carries no claim of first use here, and appears only as the consequence that outlives a cancellation. Four figures in circulation were checked and left out: the Arkes and Blumer experimental percentages, because the paper could not be opened and the secondary versions disagree; the Keil, Mann and Rai survey sample size, for the same reason; the RAND 80 per cent, because its own endnote points to a press article; and the MIT NANDA 95 per cent, because the document is a version 0.1 preliminary paper hosted on no institutional domain.

How this research works  ·  Reviewed quarterly  ·  Found an error? Tell me and it is corrected on the page.

Evidence review · SS-2026-176 · Graded against the published rubric

Cite this page

Hirji, R. (2026). Which AI investments should we stop?. The SuperSkills evidence base, SS-2026-176. https://thesuperskills.com/research/which-ai-investments-should-we-stop. Last reviewed 4 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

Which AI investments should we stop?

Stop the ones where nobody can now state what the programme was supposed to change, where the only evidence of benefit is self-reported or supplied by the vendor, and where the sponsor and the assessor are the same person. Those three conditions describe escalation rather than difficulty, and escalation is the failure mode the research actually measures. Keil, Mann and Rai found that between 30 and 40 per cent of information systems projects exhibit some degree of escalation. Note what this test does not require: a counterfactual estimate of what the programme was worth. Most organisations cannot build one, and waiting for it is itself a way of not deciding.

What is escalation of commitment?

The tendency to commit further resources to a course of action that is failing, particularly when you were responsible for choosing it. Staw demonstrated it experimentally in 1976: participants personally responsible for an earlier investment allocated an average of 11.08 million dollars to the division they had chosen, against 8.89 million where someone else had chosen it, and 13.07 million where their own earlier choice had subsequently declined. It is an established concept in organisational behaviour and is not a SuperSkills coinage.

What percentage of AI projects fail?

There is no credible non-vendor measurement of this. The most repeated figure, that more than 80 per cent of AI projects fail, is cited by RAND to a press article rather than to a study. The MIT NANDA claim that 95 per cent of organisations get zero return comes from a self-described preliminary version 0.1 document, based on 52 interviews and 153 survey responses gathered at industry conferences, that is not hosted on any MIT domain. The nearest defensible figure is from the US Census Bureau: 14.5 per cent of current business AI users do not expect to be using it in six months, which is a stated intention rather than an observed discontinuation.

How do organisations stop a failing IT programme?

Montealegre and Keil, studying the baggage handling system at Denver International Airport, describe de-escalation as a four-phase process: problem recognition, re-examination of the prior course of action, search for an alternative course of action, and implementing an exit strategy. The value of the model is that it separates recognising a problem from being permitted to act on it. Most organisations reach phase one repeatedly and never reach phase two, because the person who would have to re-examine the prior course of action is the person who chose it.

In this hub

Organisations and leadership

What a leadership team actually has to decide, and what to measure.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

This argument is one a board usually meets for the first time in the room. There is the boards and leadership version, and the full range of topics and audiences.

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.