Yes, and the mechanism is more uncomfortable than most people assume, because it does not work by making anyone worse. It works by making everyone individually better in the same direction. The clearest evidence comes from a 2024 experiment published in Science Advances: writers given AI-generated story ideas produced work rated more creative, better written and more enjoyable, with the largest gains going to the least creative writers, and the resulting stories were markedly more similar to one another than stories written unaided. Every writer was individually right to use the tool. The literature that resulted was duller. That is a social dilemma rather than a failure of the technology or the people using it, and social dilemmas are not solved by trying harder.
The answer, in one line
The evidence points that way, and the mechanism is that it makes people individually better in the same direction. Doshi and Hauser, publishing in Science Advances in 2024, gave 293 writers AI-generated story ideas judged by 600 evaluators.
Definition#
Homogenisation: the narrowing of the range of what a population produces or thinks, caused by many people drawing on the same source of suggestions, even where each person's own output improves. It is measured at the level of the group and is invisible to every individual inside it.
What Doshi and Hauser measured#
Doshi and Hauser ran the study with 293 writers producing short fiction and 600 evaluators judging it. Writers were given no AI ideas, one AI idea, or five. More AI exposure produced better-rated individual stories and greater similarity between them. The authors describe it explicitly as a social dilemma: individually beneficial, collectively narrowing.
The pattern shows up in a different form in the workplace evidence. Brynjolfsson, Li and Raymond found AI assistance raised productivity by thirty percent for the newest customer-support agents and almost nothing for the most experienced, because the system transfers the patterns of high performers to everyone else. That is a genuine gain. It is also a description of convergence: the tool works by making more people produce what the best people produce, which necessarily reduces variance.
And it is not only outputs that converge. The 2025 Microsoft Research and Carnegie Mellon survey of 319 knowledge workers found that the thinking itself shifts from generating to verifying, from solving to integrating. A person evaluating a proposed answer is exploring a much smaller space than a person generating one, and the space they are exploring was defined by the model rather than by them.
Suggestions move a person's own words, not only the suggested ones#
Doshi and Hauser show convergence in a product. Hohenstein and colleagues, publishing in Scientific Reports in 2023, show it happening inside a person, under randomisation, which makes it the strongest causal evidence assembled here.
They ran two preregistered experiments on algorithmic reply suggestions in live text chat: 438 crowdworkers in 219 pairs, then 582 in 291 pairs, with the availability of smart replies randomised separately for each partner. Smart replies accounted for 14.3 per cent of messages sent. Greater use of them by one partner led the other person to write with more positive sentiment, and the effect held when the suggested messages themselves were stripped out of the sentiment score. The second experiment manipulated the emotional tone of the suggestions, and conversation sentiment followed it.
The detail that carries the argument is a null result. Merely having suggestions on screen changed nothing; the effect came through using them. So the influence is not a matter of being nudged by what you see. It travels through the act of adopting the machine's phrasing, and it then shows up in the sentences you compose yourself. Two hundred people accepting reasonable first drafts is exactly the situation this measures.
The dialect result, and what convergence costs whom#
Convergence is usually discussed as though the cost were spread evenly, a general flattening that everybody pays a little of. Fleisig and colleagues, at EMNLP in 2024, measured who actually pays.
They put around fifty native-speaker messages in each of ten varieties of English to GPT-3.5 Turbo and GPT-4, annotated ten linguistic features per variety at a Krippendorff's alpha of 0.97, and had native speakers evaluate the replies. A Standard American English input keeps 77.9 per cent of its distinctive features in the model's reply, and Standard British English 72.2 per cent. Five of the eight minoritised varieties keep 2 to 3 per cent. Indian, Nigerian and Kenyan English fall in between at 10 to 16 per cent, and retention tracks the estimated speaker population of each variety.
Retention of 2 per cent is not a preference. It is erasure of almost every marker of how a person speaks, performed silently and returned as help. The evaluation ran alongside it: replies to non-standard varieties carried 19 per cent more stereotyping, 25 per cent more demeaning content, 9 per cent less comprehension and 15 per cent more condescension. One caution belongs with the finding, because secondary coverage regularly inverts it. The paper is often reported as showing that models exaggerate or caricature dialects. The measured default is the reverse: features are stripped out, not amplified.
The lexical study everyone cites, and the number to leave alone#
The most-quoted evidence for AI changing human language is Yakura and colleagues at the Max Planck Institute for Human Development. It is worth being careful about, because the version in circulation is not the version that now exists.
The 2024 preprint analysed roughly 280,000 YouTube transcripts from academic-institution channels and reported that words the model favours rose over the following eighteen months: delve by 48 per cent, meticulous by 40, realm by 35 and adept by 51. The 51 per cent is the figure that travelled. It belongs to one word, its outcome is the proportion of videos containing that word rather than how often anyone said it, and when the authors hand-checked fifty videos, 32 per cent showed signs of someone reading from a script.
The authors have since replaced that analysis twice. The current version, posted in July 2026, drops the YouTube corpus for 737,083 hours of conversation across 824,634 podcast episodes screened for unscripted speech, adds a synthetic-control design, and adds a preregistered experiment with 496 participants in which a brief chatbot interaction led people to adopt its words as their own, surviving a distractor task. Neither the word "adept" nor the figure 51 appears anywhere in it.
So the paper is stronger than it was and the number is weaker than its reputation. Cite the direction and drop the percentage. Two of the original limits survive intact into the new version: the outcome is still the share of episodes containing a word rather than the frequency of a speaker using it, and the paper remains unpublished after four revisions, with no journal reference and no DOI beyond the preprint server's own.
The result about proposals, not prose#
Everything above measures words. One study measures what professionals actually put forward, and it belongs here because the question on this page is about thinking. Dell'Acqua and colleagues ran a pre-registered field experiment with 776 professionals at Procter and Gamble on real product innovation problems, randomising both AI access and whether people worked alone or in pairs. Individuals with AI matched the performance of two-person teams without it, which is the headline the study is usually quoted for. Graded entry.
The finding that matters for convergence is the second one. Without AI, research and development professionals proposed more technical solutions and commercial professionals proposed more commercially oriented ones, and the split was visible in the work. With AI, both groups produced balanced proposals whatever their training. The difference a specialist brings was gone inside a single session, in a firm, on real problems.
Two cautions belong with it, and the authors supply both. The study establishes that the proposals became more alike and says nothing about whether they became better, so whether a firm loses something when its chemists and its marketers converge is left open here. And this is one company, one task type and a single session, with Procter and Gamble supporting the institute involved, which the paper discloses. What it adds to this page is a measurement the text studies cannot reach, taken upstream of the writing: the option a professional puts on the table in the first place.
One task, one form of assistance#
The Doshi and Hauser result is one creative task, short fiction, with one form of assistance. Whether homogenisation of the same magnitude occurs in domains where novelty is judged differently, such as engineering or law, is genuinely unknown. It is also possible that the effect is transitional: as models diversify and people learn to push against them, output range may recover. Nobody has measured that, and claiming either way would be guessing.
The studies added above do not close that gap either, and each falls short in a different place. Dell'Acqua measures proposals inside one firm on one kind of problem, so it shows convergence of output and not convergence of the thinking behind it. Hohenstein is causal and preregistered, and its participants were crowdworkers discussing policy with strangers over a few minutes, which is not a colleague, a client or a marriage. Fleisig measures model output and says nothing at all about how people speak. Yakura measures how people speak and cannot isolate a single cause across a fragmenting field of models. What they establish between them is that the mechanism is real at three separate points in the chain. None of them measures an organisation, over years, which is where the claim actually matters.
There is also a reasonable objection worth taking seriously. Convergence is not automatically bad. A great deal of professional work should converge, because there is a right answer and spread around it is error rather than diversity. Radiological reporting, contract drafting and safety procedure benefit from consistency. The harm falls specifically on domains where the value lies in the range of what gets tried, which is a narrower claim than "AI makes us think alike".
The word doing the work is collective#
The important word in the Doshi and Hauser finding is collective. Almost all argument about AI is conducted at the level of the individual: will it help me, will it replace me, does it make me sharper or lazier. This is the first well-designed study showing an effect that only exists in aggregate, where no individual can detect it and no individual decision causes it.
That makes it structurally similar to capability debt, and to the reason I keep returning to drift versus design. Nobody decides to narrow the range of what an organisation thinks. It happens because two hundred people each accept a reasonable first suggestion, and the suggestions come from the same place. The output is better and the portfolio is thinner, and the only level at which anyone could notice is the level at which nobody is looking.
There is a sharper version for anyone whose work depends on being distinctive. If competitors use the same models, prompted in broadly the same way, on broadly the same public information, then the strategy that emerges is broadly the same strategy. Differentiation has historically come from somewhere: proprietary information, a particular history, an odd founder, an argument nobody else was making. Convergent tooling attacks the last of those directly, and it does it while every quarterly output looks better than it did before.
The Fleisig result adds something the Doshi and Hauser framing leaves out, and it changes who should care. A social dilemma implies a cost shared by everybody in the group. A 77.9 per cent retention rate against 2 per cent is a cost with an address. The people whose way of speaking survives contact with the model lose least, and they are also the people most likely to be designing, buying and evaluating these systems. An organisation that adopts a single assistant across a workforce is not only narrowing its range. It is choosing a voice, and the choice was made somewhere else.
This is also why I think the strange question is becoming the scarce asset. Models answer well and have no questions of their own, because they have no stake in which answer is right. The first casualty of very good answers is the odd, unpromising, slightly embarrassing line of enquiry that nobody would have suggested and that occasionally turns out to matter.
Protecting range, on purpose#
Generate before you consult. Write your own list first, however bad. Once the model has spoken, its framing is in the room and the alternatives you never saw are gone. This is the individual form of Human at the Start.
Deliberately protect variance in group settings. Have people form independent views before any shared AI-assisted document exists. Silent written positions before discussion is an old technique that works for the same reason here as it always did.
Keep a source of questions that is not the machine. Read outside your field, talk to people who disagree with you, and pay attention to what irritates you. Homogenisation is produced by individually rational choices, so resisting it has to be deliberate rather than incidental.
Measure range, not just quality. If your organisation reviews AI-assisted work, ask how different this quarter's proposals are from last quarter's, and from your competitors'. Nobody tracks this, so it moves without being noticed.
Ask whose voice the tool keeps. The Fleisig result is a procurement question as much as a linguistic one. In a workforce that does not all speak Standard American English, a single assistant applied to everyone's writing removes more from some people than others, and nothing in a standard evaluation would surface it.
Related SuperSkills research#
On what survives, what stays human. On the mechanism at organisational scale, capability debt and drift versus design. On the thinking shift, AI and critical thinking. On keeping your own register, how do I keep my own voice when using AI. On practice, using AI without dependency. See should AI remember everything about me.
Key sources
- Dell'Acqua, F., Ayoubi, C., Lifshitz, H. et al. (2025). The Cybernetic Teammate. NBER Working Paper 33641; Organization Science, June 2026. Graded entry.
- Doshi, A. R. and Hauser, O. P. (2024). Generative AI enhances individual creativity but reduces the collective diversity of novel content. Science Advances, 10(28).
- Hohenstein, J., Kizilcec, R. F., DiFranzo, D., Aghajari, Z., Mieczkowski, H., Levy, K., Naaman, M., Hancock, J. and Jung, M. F. (2023). Artificial intelligence in communication impacts language and social relationships. Scientific Reports, 13, 5487.
- Fleisig, E., Smith, G., Bossi, M., Rustagi, I., Yin, X. and Klein, D. (2024). Linguistic Bias in ChatGPT: Language Models Reinforce Dialect Discrimination. Proceedings of EMNLP 2024, 13541-13564.
- Yakura, H., Lopez-Lopez, E., Brinkmann, L., de la Serna, I., Kirfel, L., Gupta, P., Soraperra, I., Eisenmann, T. F., Wulff, D. U. and Rahwan, I. Empirical evidence of Large Language Model's influence on human spoken communication. arXiv:2409.01754, v1 September 2024, v4 July 2026. Not peer-reviewed.
- Brynjolfsson, E., Li, D. and Raymond, L. (2025). Generative AI at Work. Quarterly Journal of Economics, 140(2), 889-942. Earlier version NBER Working Paper 31161.
- Lee, H.-P. et al. (2025). The Impact of Generative AI on Critical Thinking. Microsoft Research and Carnegie Mellon, CHI 2025.
About this research#
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. Findings are attributed to the studies that produced them and kept separate from the interpretation, which is the author's. The graded evidence, including what each study does not support, is in the evidence base. Last reviewed: 7 September 2026. Corrected 7 September 2026: this page's structured data cited the Yakura preprint with the six authors of its 2024 version, and now carries the ten authors of the current one. The 51 per cent figure that circulates from that paper is named here and declined.
Essay · SS-2026-040 · 3 peer-reviewed studies, 1 working paper and 1 simulation model
Hirji, R. (2026). Does AI make everyone think alike?. The SuperSkills evidence base, SS-2026-040. https://thesuperskills.com/research/does-ai-make-everyone-think-alike. Last reviewed 7 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work