Whatever the cheap thing is an input to and cannot itself supply. That is the textbook answer and it is true, which is also why it is useless in a management meeting: everybody nods and nobody can name one. The specific answer is available, and it comes from three randomised experiments that were designed to measure something else. All three found AI raising the floor much further than the ceiling. If that holds, the capability losing value fastest in a business is being unusually good at exactly the work the model does well, and the things that appreciate are the ones a model cannot hold rather than the ones it cannot do.
The answer, in one line
Whatever the cheap capability is an input to and cannot itself supply. The specific version is sharper than the general one. Three randomised experiments found the same result independently: AI raises the floor far more than the ceiling.
Three experiments, three settings, one compression#
The best-evidenced statement anyone can make about generative AI at work concerns the shape of the distribution rather than the average gain, and three independent randomised designs agree on it.
Brynjolfsson, Li and Raymond studied the staggered introduction of a conversational assistant across 5,172 customer-support agents in 133 teams at one firm, published in the Quarterly Journal of Economics in 2025. Resolutions per hour rose 15 per cent on average. The distribution underneath that average is the finding: less skilled and less experienced workers gained 30 per cent, rising to 36 per cent for the lowest skill quintile, while the most skilled saw no significant productivity change and small declines in conversation quality and customer satisfaction. The model had been fine-tuned on the firm's own past conversations with its top performers deliberately up-weighted. It transferred a portion of their behaviour to the newest agents and gave the top performers nothing.
Noy and Zhang assigned incentivised, occupation-specific writing tasks to 453 college-educated professionals in a preregistered experiment, published in Science. Average time fell by 40 per cent and rated output quality rose by 18 per cent. The distributional result is the one this page is about: "In the treatment group, initial inequalities were more than half-erased by the treatment: the correlation between first-task and second-task grades was only 0.14." Knowing how good somebody was at the first task told you almost nothing about how good they would be at the second, once the tool was in the room.
Peng, Kalliamvakou, Cihon and Demirer recruited 95 professional programmers for a randomised trial of GitHub Copilot on a standardised task. Of the 35 in each arm who completed, the treated group finished 55.8 per cent faster, with a 95 per cent confidence interval between 21 and 89 per cent. The authors record that the study "does not examine the effects of AI on code quality", which matters for any inference about value rather than speed.
One study points the other way and belongs here rather than in a footnote. METR ran a randomised trial with 16 experienced open-source developers across 246 real tasks on large, mature repositories they knew well, and measured them 19 per cent slower with AI. That is the compression finding seen from the top end: the people it did not help were the ones who already held deep, specific, hard-won context. METR withdrew the 19 per cent as a current signal in February 2026 and now believe developers are more sped up than they were, while saying their new data is only very weak evidence for the size of the change. The direction of the original finding is contested. The heterogeneity across all four studies is not.
The scarce thing was the gap, and the gap is narrowing#
Compression is a pricing event before it is a capability event. A business that charged a premium because its people were meaningfully better than the market at drafting, summarising, first-pass analysis or standard code was selling a gap. Where a tool moves the median performer most of the way to the expert on a given task, the gap on that task narrows, and so does what anyone will pay for it.
Which direction that moves the value of a whole role depends on which tasks got cheap, and there is four decades of labour data on precisely that question. Autor and Thompson, using a content-agnostic measure of task expertise across 303 US occupations from 1980 to 2018, found that automation which removed the less expert tasks from a job raised wages and reduced employment, while automation which removed the expert tasks lowered wages and increased employment. Two occupations can automate the same share of their work and move in opposite directions on pay, according to which share it was.
Their data ends in 2018 and measures nothing about generative AI, so it is a lens rather than a forecast. As a lens it produces the single most useful question a business can ask about its own roles: of the tasks this tool is taking, were they the ones that required the most or the least of our expertise? A firm whose answer is "the most" is watching its premium erode while its throughput rises, and the throughput will show up in the numbers first.
Autor states the risk in one sentence: "The risk is the devaluation of expertise." He also argues the opposite outcome is attainable, that AI could extend the reach of expert decision-making to a larger set of workers, and he is careful about the status of the claim: "My thesis is not a forecast but a claim about what is attainable." Both halves belong in a strategy discussion. The first is what happens by default.
Four complements that appreciate, and why they are the ones#
The general answer, complements appreciate and substitutes depreciate, becomes usable once you ask a narrower question: which complements can a model not supply even in principle? Capability is the wrong axis, because capability keeps moving. Standing is the right one.
Accountability. Someone has to be answerable for the output, and a model cannot be. This is not a claim about what AI can do; it is a claim about what a client, a regulator or a court will accept. As the volume of producible work rises, the scarce input becomes the person willing to put their name on it. That scarcity is turning verification ownership into a role rather than a step.
Proprietary context. The model has read what everyone has read. It has not sat in your last three board meetings, does not know which of your customers is about to leave, and cannot see the constraint nobody wrote down. The METR result is a version of this: the developers it did not help were the ones whose advantage was context the tool could not reach. Context of that kind is built by presence over time, which makes it the asset most easily destroyed by removing the roles that produced it.
The question rather than the answer. Compression applies to answering. It does not apply to deciding what is worth asking, which is upstream of every prompt and is not itself promptable. This is the argument of what makes a good question, and the position I published in "Knowledge Is No Longer Power" (2025): once knowledge is instantly and universally available, advantage moves to judgement about which questions are worth asking.
The relationship the work is delivered inside. The WORKBank audit's own early signal points here: comparing tasks by the human agency they require against the wages of their core skills, the authors find "traditionally high-wage skills like analyzing information are becoming less emphasized, while interpersonal and organizational skills are gaining more importance". They call it an early signal rather than a measured shift. Give it exactly that weight. It is also the only large dataset pointing at the same conclusion the pricing logic reaches independently.
Every figure a business currently holds about its own gain is probably wrong#
A page about value has to say something about measurement, because most organisations answering this question will answer it from a survey of their own staff. That is an unreliable estimator, and the interesting finding is that it is unreliable in both directions.
METR are the only people to have surveyed and measured the same population on the same metric, and they say so: "To our knowledge, only one study gathered survey and field experiment results on the same population and metric, Becker et al. (2025), which finds that developers overestimate productivity gains by over 40 percentage points." In the Copilot trial the error ran the other way. Participants in both arms "estimated a 35% increase in productivity, which is an underestimation compared with the 55.8% increase in their revealed productivity". Self-report was 20 points too low in one setting and more than 40 points too high in another.
METR's own hedge is the responsible position and it is rarely quoted: "it is difficult to determine the extent to which surveys overestimate productivity gains relative to experimental data". They add a distinction that belongs in any business case, that "speed measures are likely biased upwards with respect to value measures; however, value is also more opaque and harder for respondents to think about". Hours saved is the easy quantity to produce and the wrong one to buy on.
The US Census Bureau, which runs the largest firm-level AI survey there is, refuses the causal claim in its own text: "We emphasize that the analysis in this section does not seek to identify a causal link between AI use and firm performance, but rather to explore whether AI use is associated with better firm performance in general." A national statistical agency with administrative microdata will not say adoption caused performance. An organisation with a staff survey and a licence count is in no position to.
The practical implication for this question is narrow and useful. An estimate of where AI is creating value inside a business, derived from asking people how much time it saves them, is not evidence about value and cannot be used to decide which capabilities to keep. The timing of that measurement is a separate problem with its own answer, covered in how long before you know if an AI investment worked.
Sorting revenue by what it actually prices#
- Sort your revenue by whether it prices a gap or an outcome. Work sold on the basis that your people are better at producing it is exposed to compression. Work sold on the basis that you are accountable for the result is not, and the two frequently sit in the same invoice.
- Ask the Autor and Thompson question about each role. Of the tasks now being done with AI, were those the ones that took the most expertise or the least? Ask the person doing the job, because the answer is not visible from a process map.
- Protect the repetitions that build the context you are actually selling. Proprietary context is produced by people doing the work over time. An efficiency programme that removes the junior version of that work is selling next decade's advantage to fund this quarter's, which is what capability debt names.
- Stop counting hours saved as a proxy for value. Hold out a comparison group, or accept that you have a satisfaction score rather than a result. Two of the studies above exist only because somebody was willing to randomise.
- Reprice deliberately rather than by erosion. Where a gap has closed, the market will find out. A firm that reprices on its own terms keeps the relationship; a firm that waits is repriced by a client who has run the comparison first.
Where this sits in my own argument#
"Expertise-as-a-Service" (2023) argued that fractional senior roles are structurally different from part-time ones, and that what is being bought is unbiased outside judgement at an inflection point rather than cost saved, with AI pushing those roles towards deeper specialism rather than making them redundant. The compression evidence published since is the mechanism behind that prediction: as the general layer gets cheap, what survives at a premium is the part that is specific, accountable and outside.
"Actual Intelligence" (2026) put the harder half of it. A machine can produce a competent version of almost any piece in seconds, and the years of unrewarded practice behind the good version are the part that cannot be faked. Compression is what that looks like from the buyer's side. It closes the visible gap while leaving the invisible one exactly where it was, and it takes a while for anyone to notice which of the two they were paying for.
Attribution note. The expertise measure and the wage finding are Autor and Thompson's. The devaluation warning is Autor's. The productivity results belong to their authors. Complements and substitutes are basic economics and belong to nobody. Missed reps is mine. Capability debt I have used and developed since June 2025 without claiming first use, because the phrase is in independent use elsewhere.
Three experiments are a convergence and not a law#
It does not claim that compression is universal. Three experiments in customer support, professional writing and software development, on tasks chosen to be tractable, is a convergence rather than a law. Noy and Zhang say so about their own work: "We examined a limited range of occupations and tasks, in which ChatGPT may be unusually useful." Whether the same shape appears in work with long feedback loops and no clean output measure is unknown, because nobody has measured it.
It does not claim that the value of expertise is falling. Autor and Thompson's finding runs both ways, and their data predates generative AI entirely. The claim here is narrower: the value of being better than your peers at the specific tasks a model does well is falling, and that is not the same quantity as expertise.
It does not offer a benchmark, a multiplier or a return figure. The numbers in circulation for enterprise AI returns are self-reported, inconsistent about what counts as AI spend, and published without a checkable method. The most-quoted claim in this area, that around 95 per cent of enterprise generative AI pilots produce no measurable return, traces to a report whose success criterion is whether users or executives remarked on an impact, whose own limitations section calls its figures directionally accurate rather than measured, and which this research could find on no institutional domain. It goes unused here.
It does not claim the four complements are exhaustive or ranked. They are the ones that follow from the standing argument rather than from a capability argument, and a different cut of the same logic would produce a different four.
Key sources
- Brynjolfsson, E., Li, D. and Raymond, L. (2025). Generative AI at Work. Quarterly Journal of Economics, 140(2), 889-942. Article page.
- Noy, S. and Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence. Science, 381(6654), 187-192.
- Peng, S., Kalliamvakou, E., Cihon, P. and Demirer, M. (2023). The Impact of AI on Developer Productivity: Evidence from GitHub Copilot. arXiv:2302.06590.
- Autor, D. and Thompson, N. (2025). Expertise. Journal of the European Economic Association, 23(4), 1203-1271.
- Autor, D. (2024). Applying AI to Rebuild Middle Class Jobs. NBER Working Paper 32140.
- Model Evaluation and Threat Research (2025, 2026). The early-2025 developer productivity study and its February 2026 design update, with the May 2026 self-report survey.
- Bonney, K. et al. (2024). Tracking Firm Use of AI in Real Time. US Census Bureau, CES Working Paper 24-16.
- Shao, Y. et al. (2026). Future of Work with AI Agents. arXiv:2506.06576.
Related SuperSkills research#
On the firm-level version of the question, if everyone has AI, where is the advantage and how long before you know if an AI investment worked. On the individual version, staying valuable in the age of AI and will AI replace my job. On what the compression does to development, synthetic seniority, missing rungs and capability debt. On the measurement problem, usage theatre and measuring AI adoption properly. On where the gains end up, who captures the productivity gains from AI.
About this research#
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. The Brynjolfsson, Li and Raymond figures used here are the peer-reviewed ones, 5,172 agents and 15 per cent, rather than the 5,179 and 14 per cent that appear in the working paper and in most citations of it. Its distributional result is given in this estate's own graded wording rather than as a quotation, because the published abstract could be retrieved only through a bibliographic aggregator and the article body is behind a paywall. The Noy and Zhang text was read at the corresponding author's own hosted accepted manuscript, because the publisher page returned nothing to this research. No return-on-investment benchmark appears anywhere on this page, and the section on what it does not claim says why.
Evidence review · SS-2026-165 · Graded against the published rubric
Hirji, R. (2026). What becomes more valuable in a business as AI gets cheaper?. The SuperSkills evidence base, SS-2026-165. https://thesuperskills.com/research/what-becomes-more-valuable-as-ai-gets-cheaper. Last reviewed 3 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work