By compressing the distance between a new agent and a good one, from both ends. The newest staff get much faster and better. The best staff get slightly worse. And because the system was trained on the firm's own top performers, what spreads through the workforce is a copy of behaviour that used to take years to acquire, delivered in a fortnight to people who have not acquired it.
Customer support is the occupation with the strongest field evidence anywhere. It is also the one whose public discussion runs almost entirely on company press releases. This page uses the first and declines to quote the second.
The study that watched the tool being switched off#
Brynjolfsson, Li and Raymond published Generative AI at Work in the Quarterly Journal of Economics in 2025, after an earlier working-paper version. They observed the staggered deployment of a GPT-3-based conversational assistant across 5,172 customer-support agents in 133 teams at a single Fortune 500 business-process software firm, most of them working from the Philippines. Three million chats, 1.2 million of them after deployment.
The headline: resolutions per hour rose 15 per cent on average, and 15.2 per cent in the specification with agent and tenure fixed effects.
The distribution is the actual finding.
- Less skilled and less experienced workers: a 30 per cent increase in issues resolved per hour, rising to 36 per cent for the lowest skill quintile.
- The most skilled: no significant productivity change, and small but statistically significant declines in resolution rates and customer satisfaction.
- Agents with less than a year's tenure improved. Agents beyond a year showed no effect at all.
- Treated agents with two months' tenure performed as well as untreated agents with more than six months.
A note on the numbers, because this research had them wrong. The widely quoted figures of 14 per cent overall and 34 per cent for novices come from the NBER working paper. The published article says 15 and 30, with 36 for the lowest skill quintile, and the sample is 5,172 rather than 5,179. Those corrections were applied across this site on 31 August 2026. The working-paper numbers are still in wide circulation, including in places that describe themselves as citing the QJE.
What the model was actually trained on#
Most summaries of this study skip the design detail that explains its result. The assistant was not a general chatbot bolted onto a helpdesk. It was fine-tuned on the firm's own historical customer-agent conversations, labelled with outcomes, and:
Our AI firm also up-weights the value of training chats if the chat was conducted by a top performer when training the AI.
The authors list what that was meant to capture: when to ask a clarifying question, attentiveness to a customer's concern, de-escalating a tense exchange, adapting tone, explaining a complex thing simply. Tacit behaviour, deliberately harvested. Their own reading is that generative systems here "may be capable of capturing and disseminating the behaviours of the most productive agents".
So the 30 per cent measures something other than a clever machine: the firm's best agents, distilled and redistributed. That reframes almost every claim made about this study. It is evidence that expert behaviour can be copied and delivered to a novice at speed. It says nothing yet about whether the novice acquires it.
Two other effects are worth carrying. Customer sentiment improved by half a standard deviation, and requests to speak to a manager fell about 25 per cent. Attrition among agents with under six months' experience fell by roughly 10 percentage points against a 25 per cent baseline, a 40 per cent reduction, though the authors caution that this result lacks agent fixed effects and may overstate the effect. Surveyed customer satisfaction, measured by net promoter score, showed no significant difference at all, which is a useful check on the sentiment result.
The outage evidence, and the condition attached to it#
Here is what makes this study rare. The assistant sometimes broke. Technical outages interrupted the recommendations without warning, giving the authors something almost nobody else has: a natural test of what the worker retained when the tool went away.
During outages, agents with AI exposure still handled chats faster than their own pre-AI baseline, equivalent to 15 to 25 per cent declines in chat duration. And the effect grew with exposure: an outage one month after adoption showed little advantage, an outage three months in showed a clear one.
Then the condition, which is the sentence this page exists to put in front of people:
Panel C reveals that workers with high initial adherence to AI recommendations experience significant and rapid declines in chat processing times, even during outages, relative to their pre-adoption baseline. In contrast, Panel D shows no such improvement for workers who frequently deviate from AI suggestions; they see no reduction in chat times during outage periods, even after prolonged AI access.
Learning happened, and it happened only to the workers who engaged with the suggestions and watched what customers did in response. Passing the suggestion through produced output at the time and nothing afterwards. The authors' own gloss: workers learn more by actively engaging with AI suggestions and observing firsthand how customers respond.
That is the most useful finding in the whole literature on this question, and it sits buried in a robustness section. It is also, as the authors say, their noisiest: outages are rare, and the chats occurring during one may not be comparable to the chats occurring outside one. Treat it as the best available evidence for a mechanism rather than as a measured effect size.
The result nobody quotes: the best agents got slightly worse#
Experienced, high-skill agents saw no speed benefit worth naming and small declines in the quality of their conversations and in customer satisfaction. Yet the authors record that top workers increased their adherence to the recommendations over time, "even though those recommendations marginally decrease the quality of their conversations".
An expert taking advice that makes their work slightly worse, and taking more of it as time passes, is a textbook description of automation bias, observed in payroll data rather than a laboratory. It also has a consequence the authors raise themselves: with fewer original contributions from the most skilled workers, the future training data thins out, and later versions of the system have less of the thing that made this one work.
The tool captured the top performers' behaviour, distributed it, and then began eroding the source.
Where this evidence stops#
The authors are careful, and their caveats matter more than usual because this single study carries so much of the public argument about AI and work.
- One firm, one occupation, one tool. They say the findings "should not be generalized across all occupations and AI systems", and note their setting has a relatively stable product and question set. In a fast-changing environment the tool might synthesise new practice, or entrench outdated practice from historical data.
- Medium-run, partial equilibrium. Wages, labour demand and the skill mix of new hires were not observed. They flag a possible ratchet effect if performance targets are revised upward to absorb the gain.
- Headcount is a calculation, not a finding. Their back-of-the-envelope: the firm could field the same volume with 12 per cent fewer worker-hours. Whether that becomes fewer people depends on demand elasticity, which they cannot see.
- No data window is published. The rollout ran through late 2020 and early 2021. The article body gives no explicit start and end date for the sample, so none is stated here.
On employment, the nearest thing to a sector answer is Brynjolfsson, Chandar and Chen, who find employment among 22 to 25 year olds in highly AI-exposed occupations about 19 per cent below the comparison trend, concentrated where AI substitutes rather than complements, and running through reduced hiring rather than dismissal. Observational, and they explicitly rule out economy-wide displacement on current evidence.
Why this page quotes no company case study#
There is a well-known public reversal in this sector: a large fintech that replaced a substantial share of its support workforce with an assistant, published striking numbers, and later resumed hiring humans, with its chief executive saying publicly that the quality had suffered. It is the most repeated story in the field.
No figure from it appears on this page. The numbers in circulation originate in the company's own announcements and in press interviews rather than in independent measurement, and the primary company page could not be opened to read them at source. An unverified number from an interested party, about its own product decision, is not evidence, and putting it beside a peer-reviewed field experiment would suggest they are the same kind of thing.
The absence is itself worth noticing. This sector has one of the best natural experiments in labour economics and a public conversation conducted almost entirely in vendor claims. The general shape those claims describe does have support: gains concentrate on routine volume, so the residual cases are systematically harder than the average case the workforce was sized against. Staffing to the average is the error. That much can be said without a single company's figures.
What to do if you run a support operation#
- Measure adherence, not usage. The outage evidence says the workers who engaged with suggestions kept the gain and the ones who passed them through did not. Both look identical on a usage dashboard, which is the general problem this research calls usage theatre.
- Watch your best agents for drift. They are the group the study found getting slightly worse while taking more of the advice, and they are also your training data.
- Run the outage test on purpose. The most valuable evidence in the study came from the tool breaking. A scheduled unassisted period, on a sample, tells you what your people can still do.
- Size the team for the residual, not the average. Automation removes the easy volume first, so what is left is harder per case than what the headcount was set against.
- Keep a route to a person for the cases where the cost of being wrong is high. Complaints, hardship, disputes. The evidence on gains is about routine resolution, and generalising it past that is the error the public case studies were built on.
Key sources
- Brynjolfsson, E., Li, D. and Raymond, L. (2025). Generative AI at Work. Quarterly Journal of Economics, 140(2), 889-942, DOI 10.1093/qje/qjae044, Advance Access 4 February 2025. Published article.
- Brynjolfsson, E., Chandar, B. and Chen, R. (2026). Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence. Stanford Digital Economy Lab.
Every figure attributed to the QJE article was read from the published text, which differs from the working paper on three of them. Nothing on this page is taken from a company announcement.
Related SuperSkills research#
On what the productivity finding does and does not mean for development, how humans learn with AI, missed reps and synthetic seniority. On the expert result, automation bias and automation complacency. On measuring it honestly, usage theatre and how to measure adoption properly. On the labour end, entry-level jobs and who captures the gains. On other professions, law, medicine, consulting and accounting and audit.
About this research#
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. The published Quarterly Journal of Economics article was read at source and every quotation is verbatim from it. Building this page found that this site had been citing the working-paper figures rather than the published ones across seventeen pages; those were corrected on 31 August 2026 and the correction is recorded in the build register. No company case study figures appear here because none could be verified at a primary source.
Cite this
Hirji, R. (2026). How will AI change customer service? The SuperSkills Intelligence Company. Last reviewed 31 August 2026. thesuperskills.com/research/how-will-ai-change-customer-service
