← Research
Research

How will AI change customer service?

One study did what almost no other occupation has: it watched the tool being switched off. What happened next is the part worth keeping.

Last reviewed: 31 August 2026

What the 5,172-agent field study found and where it was misquoted, why the gains ran to novices and against the best performers, what the outage data shows about learning, and why the sector's most repeated case study cannot be verified at source.

Question this page answersAll 383 questions this research covers

By compressing the distance between a new agent and a good one, from both ends. The newest staff get much faster and better. The best staff get slightly worse. And because the system was trained on the firm's own top performers, what spreads through the workforce is a copy of behaviour that used to take years to acquire, delivered in a fortnight to people who have not acquired it.

Customer support is the occupation with the strongest field evidence anywhere. It is also the one whose public discussion runs almost entirely on company press releases. This page uses the first and declines to quote the second.

The study that watched the tool being switched off#

Brynjolfsson, Li and Raymond published Generative AI at Work in the Quarterly Journal of Economics in 2025, after an earlier working-paper version. They observed the staggered deployment of a GPT-3-based conversational assistant across 5,172 customer-support agents in 133 teams at a single Fortune 500 business-process software firm, most of them working from the Philippines. Three million chats, 1.2 million of them after deployment.

The headline: resolutions per hour rose 15 per cent on average, and 15.2 per cent in the specification with agent and tenure fixed effects.

The distribution is the actual finding.

A note on the numbers, because this research had them wrong. The widely quoted figures of 14 per cent overall and 34 per cent for novices come from the NBER working paper. The published article says 15 and 30, with 36 for the lowest skill quintile, and the sample is 5,172 rather than 5,179. Those corrections were applied across this site on 31 August 2026. The working-paper numbers are still in wide circulation, including in places that describe themselves as citing the QJE.

What the model was actually trained on#

Most summaries of this study skip the design detail that explains its result. The assistant was not a general chatbot bolted onto a helpdesk. It was fine-tuned on the firm's own historical customer-agent conversations, labelled with outcomes, and:

Our AI firm also up-weights the value of training chats if the chat was conducted by a top performer when training the AI.

The authors list what that was meant to capture: when to ask a clarifying question, attentiveness to a customer's concern, de-escalating a tense exchange, adapting tone, explaining a complex thing simply. Tacit behaviour, deliberately harvested. Their own reading is that generative systems here "may be capable of capturing and disseminating the behaviours of the most productive agents".

So the 30 per cent measures something other than a clever machine: the firm's best agents, distilled and redistributed. That reframes almost every claim made about this study. It is evidence that expert behaviour can be copied and delivered to a novice at speed. It says nothing yet about whether the novice acquires it.

Two other effects are worth carrying. Customer sentiment improved by half a standard deviation, and requests to speak to a manager fell about 25 per cent. Attrition among agents with under six months' experience fell by roughly 10 percentage points against a 25 per cent baseline, a 40 per cent reduction, though the authors caution that this result lacks agent fixed effects and may overstate the effect. Surveyed customer satisfaction, measured by net promoter score, showed no significant difference at all, which is a useful check on the sentiment result.

The outage evidence, and the condition attached to it#

Here is what makes this study rare. The assistant sometimes broke. Technical outages interrupted the recommendations without warning, giving the authors something almost nobody else has: a natural test of what the worker retained when the tool went away.

During outages, agents with AI exposure still handled chats faster than their own pre-AI baseline, equivalent to 15 to 25 per cent declines in chat duration. And the effect grew with exposure: an outage one month after adoption showed little advantage, an outage three months in showed a clear one.

Then the condition, which is the sentence this page exists to put in front of people:

Panel C reveals that workers with high initial adherence to AI recommendations experience significant and rapid declines in chat processing times, even during outages, relative to their pre-adoption baseline. In contrast, Panel D shows no such improvement for workers who frequently deviate from AI suggestions; they see no reduction in chat times during outage periods, even after prolonged AI access.

Learning happened, and it happened only to the workers who engaged with the suggestions and watched what customers did in response. Passing the suggestion through produced output at the time and nothing afterwards. The authors' own gloss: workers learn more by actively engaging with AI suggestions and observing firsthand how customers respond.

That is the most useful finding in the whole literature on this question, and it sits buried in a robustness section. It is also, as the authors say, their noisiest: outages are rare, and the chats occurring during one may not be comparable to the chats occurring outside one. Treat it as the best available evidence for a mechanism rather than as a measured effect size.

The result nobody quotes: the best agents got slightly worse#

Experienced, high-skill agents saw no speed benefit worth naming and small declines in the quality of their conversations and in customer satisfaction. Yet the authors record that top workers increased their adherence to the recommendations over time, "even though those recommendations marginally decrease the quality of their conversations".

An expert taking advice that makes their work slightly worse, and taking more of it as time passes, is a textbook description of automation bias, observed in payroll data rather than a laboratory. It also has a consequence the authors raise themselves: with fewer original contributions from the most skilled workers, the future training data thins out, and later versions of the system have less of the thing that made this one work.

The tool captured the top performers' behaviour, distributed it, and then began eroding the source.

Where this evidence stops#

The authors are careful, and their caveats matter more than usual because this single study carries so much of the public argument about AI and work.

On employment, the nearest thing to a sector answer is Brynjolfsson, Chandar and Chen, who find employment among 22 to 25 year olds in highly AI-exposed occupations about 19 per cent below the comparison trend, concentrated where AI substitutes rather than complements, and running through reduced hiring rather than dismissal. Observational, and they explicitly rule out economy-wide displacement on current evidence.

Why this page quotes no company case study#

There is a well-known public reversal in this sector: a large fintech that replaced a substantial share of its support workforce with an assistant, published striking numbers, and later resumed hiring humans, with its chief executive saying publicly that the quality had suffered. It is the most repeated story in the field.

No figure from it appears on this page. The numbers in circulation originate in the company's own announcements and in press interviews rather than in independent measurement, and the primary company page could not be opened to read them at source. An unverified number from an interested party, about its own product decision, is not evidence, and putting it beside a peer-reviewed field experiment would suggest they are the same kind of thing.

The absence is itself worth noticing. This sector has one of the best natural experiments in labour economics and a public conversation conducted almost entirely in vendor claims. The general shape those claims describe does have support: gains concentrate on routine volume, so the residual cases are systematically harder than the average case the workforce was sized against. Staffing to the average is the error. That much can be said without a single company's figures.

What to do if you run a support operation#

Key sources

Every figure attributed to the QJE article was read from the published text, which differs from the working paper on three of them. Nothing on this page is taken from a company announcement.

On what the productivity finding does and does not mean for development, how humans learn with AI, missed reps and synthetic seniority. On the expert result, automation bias and automation complacency. On measuring it honestly, usage theatre and how to measure adoption properly. On the labour end, entry-level jobs and who captures the gains. On other professions, law, medicine, consulting and accounting and audit.

About this research#

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. The published Quarterly Journal of Economics article was read at source and every quotation is verbatim from it. Building this page found that this site had been citing the working-paper figures rather than the published ones across seventeen pages; those were corrected on 31 August 2026 and the correction is recorded in the build register. No company case study figures appear here because none could be verified at a primary source.

How this research works  ·  Reviewed quarterly  ·  Found an error? Tell me and it is corrected on the page.

Cite this

Hirji, R. (2026). How will AI change customer service? The SuperSkills Intelligence Company. Last reviewed 31 August 2026. thesuperskills.com/research/how-will-ai-change-customer-service

In this hub

Professions and sectors

Where the pressure lands first, profession by profession.

The work

Where the writing comes from.

These essays draw on research across more than 200 organisations in 30 countries. See the wider body of work, or bring it into your organisation.

All research →
Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory and coaching  ·  Enquire