← Research
Research

How long should we give an AI investment before deciding whether it worked?

Three different questions are hiding inside this one, and they run on three different clocks. Reviewing all of them at month twelve produces a confident answer to none.

Last reviewed: 2 September 2026

Why the J-curve makes an early review misleading in a predictable direction, what two years of Danish payroll records actually showed, why the fastest available figures are self-reported, and a staged review a board can defend.

Question this page answersQuestion this page partly answersAll 811 questions this research covers

Longer than the budget cycle that funded it, and the reason is a measurement artefact rather than patience. Brynjolfsson, Rock and Syverson showed that general purpose technologies require intangible complements that national accounts capture badly, so measured productivity is understated early and overstated later. Their correction put US total factor productivity 15.9 per cent above official measures by the end of 2017. Humlum and Vestergaard, looking at roughly 25,000 Danish workers two years after ChatGPT, found precise nulls on earnings and hours alongside heavy task reorganisation. A defensible review is therefore staged: reorganisation at six to twelve months, quality and error rates at twelve to twenty-four, and unit economics no earlier than two years. Anything faster is measuring enthusiasm.

The answer, in one line

Longer than one budget cycle, and in stages rather than at a single date. Reorganisation of the work is visible at six to twelve months. Quality and error rates against a pre-deployment baseline take twelve to twenty-four months.

Share as a card

Three clocks, and most business cases wind only one#

The question is unanswerable as a single number because three different things are being asked about, and they move at different speeds.

A business case that promises hours saved in year one and then reports hours saved in year one has answered none of the three. It has confirmed that the first clock is running.

The J-curve says early measurement misleads in a known direction#

Brynjolfsson, Rock and Syverson published the mechanism in the American Economic Journal: Macroeconomics in January 2021. A general purpose technology enables and requires significant complementary investments, and those investments are largely intangible: retraining, process redesign, data work, new controls, the accumulated knowledge of which tasks to hand over. National accounts do not measure them well. The result is a curve. Early on, the intangible investment is a real cost that shows up in the accounts while its output does not, so measured productivity growth is understated. Later, when the intangibles are harvested, measured growth is overstated because the earlier investment was never booked.

Applying their method to US data on computer hardware and software, they found the intangible-adjusted TFP level 15.9 per cent higher than official measures by the end of 2017. For a technology diffusing in the 1990s and 2000s, that is the size of the gap between what the statistics said and what had actually happened.

The practical consequence for a chief financial officer is uncomfortable. A programme reviewed at month twelve will look worse than it is, and a programme reviewed at month forty-eight may look better than it is, and both errors have the same cause. The only reading that survives both is one that tracks the intangible investment directly: how many processes were genuinely redesigned rather than accelerated, how many people were retrained to a checkable standard, how much of the data work was finished.

Two years of Danish administrative data, and no movement in pay#

Humlum and Vestergaard linked adoption surveys to administrative labour records for roughly 25,000 workers across 7,000 Danish workplaces in eleven exposed occupations. Two years after ChatGPT, they found precise null effects on earnings and hours, ruling out effects larger than 2 per cent. Precise nulls are stronger than an absence of findings: the confidence intervals are tight enough to exclude a large effect rather than merely failing to detect one.

Underneath the nulls, the work had moved. The same study reports substantial task reorganisation and the appearance of new tasks in AI oversight and integration. This is the sequence a J-curve predicts, observed in one small, high-trust, high-wage, heavily unionised economy, two years in. The authors are explicit that Denmark may not generalise and that two years is early. Both caveats are load-bearing and both are usually dropped when the study is quoted.

Read against the investment question, it says something specific. If a firm's own two-year read shows reorganised work and unchanged unit costs, that is the expected pattern rather than a failure. If a firm's two-year read shows unchanged work and improved reported costs, something is being measured badly.

The fastest figures are the ones that cannot carry the weight#

The UK Government Digital Service ran the largest Microsoft 365 Copilot deployment anywhere, 20,000 licences across twelve organisations between 30 September and 31 December 2024. The headline of 26 minutes saved a day was computed from the midpoints of tick-box bands, with the largest savings estimated at 60 minutes, and the report's own conclusions state it was not possible to identify how the saved time was spent. The Department for Work and Pensions evaluated 3,549 staff and produced 19 minutes by regression, then published a limitations chapter naming no baseline, first-come first-served licence allocation, self-selection towards enthusiasts which it says may lead to overestimation, non-response bias and acquiescence bias.

Both are transparent, careful public evaluations. Neither timed anyone. They are the best available version of a self-reported saving, and the best available version is still a self-reported saving.

METR's randomised study puts a number on how far that can go wrong. Sixteen experienced open-source developers, 246 real tasks on repositories they knew well, randomly assigned to permit or prohibit AI tools. Measured result: 19 per cent slower with the tools. They had forecast a 24 per cent speed-up, and after finishing the tasks and experiencing the slowdown they still estimated AI had made them about 20 per cent faster. METR themselves withdrew the 19 per cent as a current signal on 24 February 2026, after a second study design produced different raw numbers and heavy self-selection in task submission. The durable finding is the perception gap, not the size. A forty-point error in the wrong direction, held after the fact, is what an hours-saved survey is exposed to.

Adoption is not a result, and the surveys measuring it are 60 points apart#

Jeffrey Allen at the Federal Reserve Board set three US measures side by side in April 2026. The Census Business Trends and Outlook Survey reports about 18 per cent of firms at the end of 2025. The Atlanta Fed's Survey of Business Uncertainty, which asks senior leaders, gives an employment-weighted 78 per cent. The Real-Time Population Survey, which asks individuals, gives about 41 per cent of the workforce using generative AI at work. Allen's guidance on which to use is worth quoting exactly, because it is the discipline most internal dashboards lack:

The BTOS is the best source for an estimate of the percentage of U.S. businesses that have adopted AI. The RPS is the best source for an estimate of the share of the labor force that uses GenAI at work. Finally, the SBU, which estimates the share of the labor force working at firms that have adopted AI, is a good upper bound on the scope of access to AI tools at work.

He also records that the Census Bureau broadened its question in November 2025, from use "in producing goods or services" to use "in any of its business functions", and that the "do not know" rate ran at 10 to 11 per cent of respondents. An organisation running its own adoption tracker has all the same problems and none of the methodological documentation. See how to measure AI adoption properly and usage theatre.

The intensive margin is where the interesting number hides#

Allen reports that daily generative AI use at work stood at 12 per cent in November 2025 against 40.7 per cent for any use, and that the Survey of Business Uncertainty found a plurality of users, 35 per cent, using AI up to one hour a week, with 29 per cent using it between one and five hours. The Census diffusion paper found 65 per cent of firms limit worker task use to three or fewer tasks.

A programme where most licence holders touch the tool for under an hour a week has not yet had the opportunity to produce a measurable financial effect, whatever the adoption dashboard says. That is not a reason to cancel it at month twelve. It is a reason to stop expecting a month-twelve answer, and to measure depth of use rather than breadth of licence in the meantime.

A staged review that can survive a board challenge#

Rahim's earlier position on the hours-saved metric#

"Time-as-a-Service" (2023) argued that the emerging business model would not be software but time, that firms would monetise the time they gave back, and that the industries ripest for disruption were the ones with the longest waits. It was written two years before hours saved became the standard unit of AI business cases. "Rules Before Tools" (2025) named what fills the gap in the meantime, including evaluation debt: the accumulating cost of deploying faster than you can assess. "You're not adopting AI. You're paying for it." (Irish Tech News, 22 July 2026) is the shortest statement of the distinction between procurement and adoption.

Attribution note. Evaluation debt, the integration tax and shadow-AI amnesty are Hirji's terms from that 2025 essay. The J-curve is Brynjolfsson, Rock and Syverson's. The unclaimed hour appears on this estate as a description and carries no claim of first use.

No payback figure here, and no dataset to take one from#

It does not claim a universal payback period. The staged gates above are a reasoning structure derived from the J-curve argument and the Danish timing, not an estimate from a dataset of AI programmes. No such dataset exists in public with a credible counterfactual.

It does not claim the Danish nulls will hold. Two years is early, Denmark is small and unusual, and the authors say so. Nor does it claim Acemoglu's 0.66 per cent is the right number: his model works through task-level cost savings and would not capture effects running through new products, new tasks or capability change, which he states.

It does not claim the GDS and DWP evaluations are wrong. They are unusually honest about their own limits, and that honesty is the reason they are used here. What they cannot do is establish a time saving, because neither measured time.

Key sources

On measurement, how to measure AI adoption properly, usage theatre and the most quoted AI statistics, checked. On the strategy question, where advantage comes from when everyone has AI and who captures the productivity gains. On stopping, deployment is not a ratchet. On the hidden cost, capability debt and the unclaimed hour.

About this research#

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. The J-curve abstract and the Federal Reserve note were read at source for this page and every figure taken from them was checked against the document. Vendor and consultancy return-on-investment benchmarks are deliberately absent: the ones in circulation are self-reported, definitionally inconsistent about what counts as AI spend, and published without a method that can be checked.

How this research works  ·  Reviewed quarterly  ·  Found an error? Tell me and it is corrected on the page.

Evidence review · SS-2026-160 · Graded against the published rubric

Cite this page

Hirji, R. (2026). How long should we give an AI investment before deciding whether it worked?. The SuperSkills evidence base, SS-2026-160. https://thesuperskills.com/research/how-long-before-you-know-if-an-ai-investment-worked. Last reviewed 2 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

How long should we give an AI investment before deciding whether it worked?

Longer than one budget cycle, and in stages rather than at a single date. Reorganisation of the work is visible at six to twelve months. Quality and error rates against a pre-deployment baseline take twelve to twenty-four months. Unit economics with an explicit counterfactual should not be judged before two years. The reason is the productivity J-curve of Brynjolfsson, Rock and Syverson: a general purpose technology requires intangible complements that national accounts capture badly, so measured productivity is understated early and overstated later.

What is the productivity J-curve?

A pattern identified by Erik Brynjolfsson, Daniel Rock and Chad Syverson in the American Economic Journal: Macroeconomics in January 2021. General purpose technologies enable and require significant complementary investments that are often intangible and poorly measured in national accounts, which leads to underestimation of productivity growth in the technology's early years and overestimation later, when the benefits of those intangible investments are harvested. Applying their method to US data on computer hardware and software, they found a total factor productivity level 15.9 per cent higher than official measures by the end of 2017.

Why can't we just measure hours saved?

Because the available hours-saved figures are self-reported and the self-report can be wrong in the opposite direction to the truth. METR randomised 246 real coding tasks across 16 experienced developers and measured them 19 per cent slower with AI tools; the same developers had forecast a 24 per cent speed-up and still estimated a 20 per cent speed-up afterwards. The UK Government Digital Service's 26-minute figure was calculated from tick-box midpoints with the top band capped at 60 minutes, and its own report states it could not identify how the saved time was spent.

Does a null result at two years mean the investment failed?

Not on its own. Humlum and Vestergaard found precise null effects on earnings and hours, ruling out effects larger than 2 per cent, across roughly 25,000 workers in 7,000 Danish workplaces two years after ChatGPT, while also finding substantial task reorganisation and new tasks in AI oversight and integration. That sequence, work changing before money moves, is what the J-curve predicts. The result that should worry a board is the reverse: unchanged work with improved reported costs, which usually means the measurement is wrong.

In this hub

Organisations and leadership

What a leadership team actually has to decide, and what to measure.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

This argument is one a board usually meets for the first time in the room. There is the boards and leadership version, and the full range of topics and audiences.

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.