Longer than the budget cycle that funded it, and the reason is a measurement artefact rather than patience. Brynjolfsson, Rock and Syverson showed that general purpose technologies require intangible complements that national accounts capture badly, so measured productivity is understated early and overstated later. Their correction put US total factor productivity 15.9 per cent above official measures by the end of 2017. Humlum and Vestergaard, looking at roughly 25,000 Danish workers two years after ChatGPT, found precise nulls on earnings and hours alongside heavy task reorganisation. A defensible review is therefore staged: reorganisation at six to twelve months, quality and error rates at twelve to twenty-four, and unit economics no earlier than two years. Anything faster is measuring enthusiasm.
The answer, in one line
Longer than one budget cycle, and in stages rather than at a single date. Reorganisation of the work is visible at six to twelve months. Quality and error rates against a pre-deployment baseline take twelve to twenty-four months.
Three clocks, and most business cases wind only one#
The question is unanswerable as a single number because three different things are being asked about, and they move at different speeds.
- Did the work change? Fastest, and visible within a quarter or two. Who does what, which steps disappeared, which new steps appeared. Humlum and Vestergaard found substantial task reorganisation and new tasks in AI oversight and integration well before anything showed up in pay.
- Did the output get better or worse? Slower, because it requires a quality measure that existed before the deployment. Most organisations discover at this point that they never had one, and the deployment has already changed the baseline.
- Did the economics change? Slowest, and confounded by everything else the business did. Acemoglu's task-based model puts total factor productivity gains at no more than 0.66 per cent over ten years, revised to under 0.53 per cent once hard-to-learn tasks are accounted for, which sets the scale of what a macro-level signal even looks like.
A business case that promises hours saved in year one and then reports hours saved in year one has answered none of the three. It has confirmed that the first clock is running.
The J-curve says early measurement misleads in a known direction#
Brynjolfsson, Rock and Syverson published the mechanism in the American Economic Journal: Macroeconomics in January 2021. A general purpose technology enables and requires significant complementary investments, and those investments are largely intangible: retraining, process redesign, data work, new controls, the accumulated knowledge of which tasks to hand over. National accounts do not measure them well. The result is a curve. Early on, the intangible investment is a real cost that shows up in the accounts while its output does not, so measured productivity growth is understated. Later, when the intangibles are harvested, measured growth is overstated because the earlier investment was never booked.
Applying their method to US data on computer hardware and software, they found the intangible-adjusted TFP level 15.9 per cent higher than official measures by the end of 2017. For a technology diffusing in the 1990s and 2000s, that is the size of the gap between what the statistics said and what had actually happened.
The practical consequence for a chief financial officer is uncomfortable. A programme reviewed at month twelve will look worse than it is, and a programme reviewed at month forty-eight may look better than it is, and both errors have the same cause. The only reading that survives both is one that tracks the intangible investment directly: how many processes were genuinely redesigned rather than accelerated, how many people were retrained to a checkable standard, how much of the data work was finished.
Two years of Danish administrative data, and no movement in pay#
Humlum and Vestergaard linked adoption surveys to administrative labour records for roughly 25,000 workers across 7,000 Danish workplaces in eleven exposed occupations. Two years after ChatGPT, they found precise null effects on earnings and hours, ruling out effects larger than 2 per cent. Precise nulls are stronger than an absence of findings: the confidence intervals are tight enough to exclude a large effect rather than merely failing to detect one.
Underneath the nulls, the work had moved. The same study reports substantial task reorganisation and the appearance of new tasks in AI oversight and integration. This is the sequence a J-curve predicts, observed in one small, high-trust, high-wage, heavily unionised economy, two years in. The authors are explicit that Denmark may not generalise and that two years is early. Both caveats are load-bearing and both are usually dropped when the study is quoted.
Read against the investment question, it says something specific. If a firm's own two-year read shows reorganised work and unchanged unit costs, that is the expected pattern rather than a failure. If a firm's two-year read shows unchanged work and improved reported costs, something is being measured badly.
The fastest figures are the ones that cannot carry the weight#
The UK Government Digital Service ran the largest Microsoft 365 Copilot deployment anywhere, 20,000 licences across twelve organisations between 30 September and 31 December 2024. The headline of 26 minutes saved a day was computed from the midpoints of tick-box bands, with the largest savings estimated at 60 minutes, and the report's own conclusions state it was not possible to identify how the saved time was spent. The Department for Work and Pensions evaluated 3,549 staff and produced 19 minutes by regression, then published a limitations chapter naming no baseline, first-come first-served licence allocation, self-selection towards enthusiasts which it says may lead to overestimation, non-response bias and acquiescence bias.
Both are transparent, careful public evaluations. Neither timed anyone. They are the best available version of a self-reported saving, and the best available version is still a self-reported saving.
METR's randomised study puts a number on how far that can go wrong. Sixteen experienced open-source developers, 246 real tasks on repositories they knew well, randomly assigned to permit or prohibit AI tools. Measured result: 19 per cent slower with the tools. They had forecast a 24 per cent speed-up, and after finishing the tasks and experiencing the slowdown they still estimated AI had made them about 20 per cent faster. METR themselves withdrew the 19 per cent as a current signal on 24 February 2026, after a second study design produced different raw numbers and heavy self-selection in task submission. The durable finding is the perception gap, not the size. A forty-point error in the wrong direction, held after the fact, is what an hours-saved survey is exposed to.
Adoption is not a result, and the surveys measuring it are 60 points apart#
Jeffrey Allen at the Federal Reserve Board set three US measures side by side in April 2026. The Census Business Trends and Outlook Survey reports about 18 per cent of firms at the end of 2025. The Atlanta Fed's Survey of Business Uncertainty, which asks senior leaders, gives an employment-weighted 78 per cent. The Real-Time Population Survey, which asks individuals, gives about 41 per cent of the workforce using generative AI at work. Allen's guidance on which to use is worth quoting exactly, because it is the discipline most internal dashboards lack:
The BTOS is the best source for an estimate of the percentage of U.S. businesses that have adopted AI. The RPS is the best source for an estimate of the share of the labor force that uses GenAI at work. Finally, the SBU, which estimates the share of the labor force working at firms that have adopted AI, is a good upper bound on the scope of access to AI tools at work.
He also records that the Census Bureau broadened its question in November 2025, from use "in producing goods or services" to use "in any of its business functions", and that the "do not know" rate ran at 10 to 11 per cent of respondents. An organisation running its own adoption tracker has all the same problems and none of the methodological documentation. See how to measure AI adoption properly and usage theatre.
The intensive margin is where the interesting number hides#
Allen reports that daily generative AI use at work stood at 12 per cent in November 2025 against 40.7 per cent for any use, and that the Survey of Business Uncertainty found a plurality of users, 35 per cent, using AI up to one hour a week, with 29 per cent using it between one and five hours. The Census diffusion paper found 65 per cent of firms limit worker task use to three or fewer tasks.
A programme where most licence holders touch the tool for under an hour a week has not yet had the opportunity to produce a measurable financial effect, whatever the adoption dashboard says. That is not a reason to cancel it at month twelve. It is a reason to stop expecting a month-twelve answer, and to measure depth of use rather than breadth of licence in the meantime.
A staged review that can survive a board challenge#
- Before deployment, record the baseline you will be judged against. Cycle time, error and rework rate, quality sample, cost per unit of work. Once the tool is in, the baseline is gone and no amount of later analysis recovers it. This single step is the difference between an evaluation and an anecdote.
- Month six to twelve, review reorganisation only. Which tasks moved, which are new, who now spends time on oversight and integration. Do not ask for a financial number at this gate, and do not accept one.
- Month twelve to twenty-four, review quality and error rates against the baseline. This is the gate that catches the expensive failure mode, where throughput rises and defect rates rise with it, invisibly, until a customer or a regulator finds them.
- Month twenty-four onwards, review unit economics. With an explicit counterfactual: what would this cost have been without the programme, given everything else that changed.
- At every gate, track the intangible investment separately from the licence spend. Retraining completed, processes redesigned, data work finished, controls written. The J-curve argument says this line is the leading indicator and the licence line is not.
- At every gate, ask what capability has been lost. If the people who used to do the work can no longer do it unaided, a cost saving has been converted into capability debt and the bill arrives later. See the capability audit.
- Write the stopping condition before you start. Not a review date, a threshold: what result would cause this to be stopped. Deployment is not a ratchet covers why this is the clause organisations skip.
Rahim's earlier position on the hours-saved metric#
"Time-as-a-Service" (2023) argued that the emerging business model would not be software but time, that firms would monetise the time they gave back, and that the industries ripest for disruption were the ones with the longest waits. It was written two years before hours saved became the standard unit of AI business cases. "Rules Before Tools" (2025) named what fills the gap in the meantime, including evaluation debt: the accumulating cost of deploying faster than you can assess. "You're not adopting AI. You're paying for it." (Irish Tech News, 22 July 2026) is the shortest statement of the distinction between procurement and adoption.
Attribution note. Evaluation debt, the integration tax and shadow-AI amnesty are Hirji's terms from that 2025 essay. The J-curve is Brynjolfsson, Rock and Syverson's. The unclaimed hour appears on this estate as a description and carries no claim of first use.
No payback figure here, and no dataset to take one from#
It does not claim a universal payback period. The staged gates above are a reasoning structure derived from the J-curve argument and the Danish timing, not an estimate from a dataset of AI programmes. No such dataset exists in public with a credible counterfactual.
It does not claim the Danish nulls will hold. Two years is early, Denmark is small and unusual, and the authors say so. Nor does it claim Acemoglu's 0.66 per cent is the right number: his model works through task-level cost savings and would not capture effects running through new products, new tasks or capability change, which he states.
It does not claim the GDS and DWP evaluations are wrong. They are unusually honest about their own limits, and that honesty is the reason they are used here. What they cannot do is establish a time saving, because neither measured time.
Key sources
- Brynjolfsson, E., Rock, D. and Syverson, C. (2021). The Productivity J-Curve: How Intangibles Complement General Purpose Technologies. American Economic Journal: Macroeconomics, 13(1), 333-72.
- Humlum, A. and Vestergaard, E. (2025). Still Waters, Rapid Currents: Early Labor Market Transformation under Generative AI.
- Acemoglu, D. (2024). The Simple Macroeconomics of AI.
- Allen, J. S. (2026). Monitoring AI Adoption in the U.S. Economy. FEDS Notes, 3 April 2026.
- Bonney, K. et al. (2026). The Microstructure of AI Diffusion. US Census Bureau, CES-26-25.
- Becker, J. et al., METR (2025 and 2026). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity and the February 2026 update.
- Government Digital Service (2025). Microsoft 365 Copilot Experiment: Cross-Government Findings Report.
- Department for Work and Pensions (2026). Microsoft 365 Copilot evaluation.
Related SuperSkills research#
On measurement, how to measure AI adoption properly, usage theatre and the most quoted AI statistics, checked. On the strategy question, where advantage comes from when everyone has AI and who captures the productivity gains. On stopping, deployment is not a ratchet. On the hidden cost, capability debt and the unclaimed hour.
About this research#
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. The J-curve abstract and the Federal Reserve note were read at source for this page and every figure taken from them was checked against the document. Vendor and consultancy return-on-investment benchmarks are deliberately absent: the ones in circulation are self-reported, definitionally inconsistent about what counts as AI spend, and published without a method that can be checked.
Evidence review · SS-2026-160 · Graded against the published rubric
Hirji, R. (2026). How long should we give an AI investment before deciding whether it worked?. The SuperSkills evidence base, SS-2026-160. https://thesuperskills.com/research/how-long-before-you-know-if-an-ai-investment-worked. Last reviewed 2 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work