- How do you run a meeting when AI attends it?
- Does AI increase managerial surveillance?
- Does AI reduce workers' autonomy?
- Does AI increase the number of decisions each person has to make?
- Are executives reading summaries instead of the source?
- How do you brief a team when AI drafts everything?
- Does AI change team decision-making?
- Does AI increase organisational sycophancy?
Almost every study of AI at work measures an individual. One person, one task, faster or slower, better or worse. Teams are not collections of individuals, though, and the few studies that look at the group find something the individual studies cannot: everybody gets better and the room gets narrower.
Everyone improves. The room converges.#
Doshi and Hauser ran an online experiment with 293 writers producing short fiction, assessed by 600 evaluators. Stories written with AI assistance were rated more creative, better written and more enjoyable, and the largest gains went to the least creative writers. They were also markedly more similar to one another.
That is a social dilemma rather than a defect. Every writer is individually right to use it. The body of work gets duller anyway. And nobody inside the experiment could have spotted it, because each person’s own output genuinely improved.
The same shape appears in a corporate setting. Dell’Acqua and colleagues ran a pre-registered field experiment with 776 professionals at Procter and Gamble on real product innovation problems, randomising both AI access and whether people worked alone or in pairs. Two results matter here. Individuals with AI matched the performance of two-person teams without it. And AI removed the functional split: without it, research and development professionals proposed more technical solutions while commercial professionals proposed more commercial ones, whereas professionals using AI produced balanced solutions regardless of their background.
Read that second finding as a manager rather than as a researcher. The reason you put an engineer and a marketer in the same room is that they see different things. If both arrive having consulted the same model, some of that difference has already been averaged away before anybody speaks. One firm, one task type, a single session, and Procter and Gamble part-funded the institute involved, making this a signal rather than a settled result. It points the same way as the writing study, from a completely different setting.
The tool agrees with you far more than a colleague would#
If a team loses range, the obvious repair is disagreement. Which makes the sycophancy evidence the most operationally important material on this page.
Cheng and colleagues tested eleven models against human responses on interpersonal advice, then ran two preregistered experiments with 1,604 participants, including live interaction on a real personal conflict. Models affirmed users’ actions about 50 per cent more often than humans did, including 47 per cent endorsement on prompts describing clearly harmful behaviour. Interacting with a sycophantic model reduced participants’ willingness to repair the conflict and increased their conviction that they were in the right. Participants rated the sycophantic model as higher quality.
Note the limit before drawing the conclusion: those scenarios are interpersonal advice rather than technical or analytical judgement, and the effect on a factual decision has not been shown. What the study does establish is that agreement changes what people subsequently do, and that preference runs in the opposite direction from benefit.
Sharma and colleagues explain why this is structural rather than a fault in one product. Across five production assistants and the human preference datasets used to train them, all five exhibited sycophancy consistently, and both humans and the preference models trained on their judgements preferred convincingly written sycophantic responses over correct ones a non-negligible share of the time. Optimising against those preferences sometimes sacrificed truthfulness. Agreeableness will not be patched out. Training on human approval produces it.
So the organisational question is not whether models flatter people. It is how much of your challenge function used to come from colleagues, and how much of it has quietly moved to a system that affirms half again as often as a human would. Every draft checked with a model instead of a sceptical peer is a small transfer in that direction, and nobody logs it.
Where team decision-making actually loses#
Vaccaro, Almaatouq and Malone’s preregistered meta-analysis of 106 experimental studies and 370 effect sizes found human and AI combinations performed significantly worse on average than the better of human or AI alone, at Hedges’ g of -0.23. The detail that matters for teams: the losses were concentrated in decision-making and the gains in content creation.
Which gives a usable rule. A team using AI to produce, draft, explore or summarise is operating where the evidence is favourable. A team using it to decide is operating where the evidence says undesigned pairing subtracts. Most teams do both with the same tool, in the same session, without marking the transition.
Does it increase the number of decisions each person makes?#
Probably, and the argument for it is old. Bainbridge’s Ironies of Automation, from 1983, observed that automating the routine parts of a task leaves the human with the hardest residue, monitoring and exception handling, while removing the routine practice that built the competence to handle it. Automation makes the remaining human role harder rather than easier.
Applied to a team, drafting moves to the model and reviewing moves to the person, so the volume of small accept-or-reject judgements rises while the practice that made those judgements reliable falls. This is a 1983 process-control argument extended by analogy rather than a measurement of generative AI, and it should be held that loosely. It is also the most consistent thing people report when asked what actually changed.
Autonomy, surveillance, and what is not known#
These two get asserted more confidently than the evidence supports, so it is worth separating what has been observed from what is inferred.
On autonomy there is one useful study. LaborIA, a French Ministry of Labour project with Inria and Matrice, surveyed 250 decision-makers in firms of more than 50 staff, ran longitudinal interviews and six ethnographic field sites. It names a conflit de rationalité, a clash of rationalities: managers justify AI by error reduction at 81 per cent, performance at 75 per cent and removing drudgery at 74 per cent, while the fieldwork shows workers becoming the system’s de facto trainers. Management and staff inside the same organisation describe different events. The qualitative core rests on six sites and ten repeated interviews, so treat it as a well-observed hypothesis rather than a measured effect.
On surveillance, this estate holds no direct evidence either way, and neither does most of the literature being cited for it. What can be said is that the infrastructure is arriving for other reasons: Article 12 of the EU AI Act requires high-risk systems to technically allow automatic recording of events across the system lifetime, for risk identification and post-market monitoring. Logging built for conformity is logging that exists. Whether organisations turn it on people is a choice nobody has measured yet.
Briefing a team when the first draft is always machine-made#
Four practical changes follow from the evidence above rather than from preference.
- Brief for divergence before anybody opens a model. The homogenisation in both studies happens at the input. Once everyone has consulted the same system, asking for different perspectives in the meeting is asking for something that has already been averaged.
- Ask who has not used it on this. A single unassisted view is now a scarce input rather than a slower one, and nothing else reliably shows whether the room has converged.
- Mark the transition from producing to deciding. Say it out loud when a session moves from drafting to choosing, because the evidence changes sign at that boundary.
- Protect the disagreement you have left. Edmondson’s field study of 51 teams found that psychological safety predicts learning behaviour, and that learning behaviour is what carries safety through to performance. If the questions people used to ask each other now go to a model that affirms half again as often, the mechanism is being drained rather than the meeting being shortened.
Meetings, and executives reading summaries#
Both belong to the same question: what happens when the thing a group reasons over is an artefact nobody in the room produced.
The team-level risk in a summarised meeting is not accuracy, which is usually adequate. The risk is that a summary is a set of decisions about what mattered, made by a system with no stake in the outcome and no knowledge of who in the room was quietly unconvinced. Dissent that was expressed weakly, which is how most dissent in organisations is expressed, does not survive compression. For the individual version of both questions, see should AI attend my meetings and should I let AI summarise everything I read.
On executives specifically, there is no measurement of how often senior people now read a summary rather than a source, and anybody offering you a figure has estimated it. The structural point stands without one: as summaries move up a hierarchy, each layer compresses, and the person with the most authority to act is furthest from the material. That was true of human briefing notes too. What has changed is the cost of producing one, and therefore the number of layers that now have them.
Where this sits in my own argument#
My central claim is that AI comes for judgement before it comes for jobs, and this is the clearest team-level version of it I have found. Nobody decides to narrow the range of views in a room. It narrows because every individual made a sensible choice, which is exactly the shape of drift: an outcome nobody chose, arrived at by people all doing something defensible.
Two of the seven capabilities in SuperSkills sit directly on this. Curiosity is what generates a view the model did not supply, and empathy is what makes somebody willing to say it in a room that has already converged. Both get harder to practise as the challenge function moves to a system that agrees.
What I have observed in organisations#
The pattern I now see most often in meetings is a general AI consensus. People arrive having read a summary rather than the document, having asked a model to prepare them, and the room converges quickly on a position nobody had to argue for. It is worst in fast-moving environments, where the speed is the whole excuse.
In creative teams it is visible in the work itself, where the creative has started to look the same. It is not confined to creatives. What runs across all of it is an assumption that the reading has been done for you: people do not open the document, because the summary is right there and takes a minute.
This is the one place where my own observation runs ahead of the published evidence. Nobody has measured how often senior people now read a summary instead of a source. I have watched it become normal.
What this page does not claim#
It does not claim teams get worse. Doshi and Hauser found better individual output. Dell’Acqua found individuals with AI matching pairs without it. Both are real gains, and an organisation may reasonably decide the convergence is worth paying for.
It does not claim these findings generalise. One is short fiction, one is a single innovation exercise at one firm with that firm’s financial support, and the sycophancy experiments concern interpersonal advice rather than analytical judgement. Each names its own limits and those limits are repeated here rather than dropped.
And it does not claim the convergence is deliberate or that anybody is at fault. The mechanism is the opposite of a bad decision: it is a great many good individual decisions producing an outcome nobody chose and nobody can see from where they are standing.
Cite this
Hirji, R. (2026). What AI does to a team. The SuperSkills Intelligence Company. Last reviewed 1 September 2026. thesuperskills.com/research/what-ai-does-to-a-team