← Research
Research · Question

Can AI stand in for real people in a survey?

What Pew did, how far off the AI respondents were and in which direction, why two models gave two publics, where Pew draws its own line, and the rule for any figure about people that reaches a board.

Last reviewed: 2 October 2026 · Next review due: 2 October 2027

Not yet, on the most careful test published so far. Pew's AI twins of its own panel members missed the real answers by an average of 12 points and erased minority answers. A simulated respondent is a model's guess about people who were not asked. An evidence review by Rahim Hirji; every figure resolves to a graded entry in the evidence base that says what it does not show.

Questions this page answersAll 1156 questions this research covers

Not yet, on the most careful test published so far. Pew Research Center reported on 30 September 2026 that AI respondents built as digital twins of its own panel members missed the real answers by an average of 12 percentage points across nearly 300 questions, and by more than 15 points on about 28 per cent of them. The errors had a direction. The synthetic public knew more than the real one, agreed with itself more, and on nearly half the questions left at least one answer unchosen. Pew concluded that AI polling does not replace surveying real people. The finding reaches further than polling. A simulated customer or employee is a model’s guess about people who were not asked, and a decision taken on it has been handed to the model without anyone saying so.

The answer, in one line

Not yet. Pew Research Center reported on 30 September 2026 that AI respondents modelled on its own panel members differed from the real answers by an average of 12 percentage points across nearly 300 questions, and by more than 15 points on about 28 per cent of them.

Share as a card

What Pew did#

The report, by Athena Chapekis and five colleagues, gave the idea its best chance. Each AI respondent was a twin of a real member of Pew’s American Trends Panel: “We gave the model a wide range of information about each panelist, including their self-reported demographic information and their responses to questions on our 2025 political typology survey.” The twins then took three of the panel’s surveys from the first half of 2026, under the same instructions and survey programming as the people who had taken them. All results unless noted come from Anthropic’s Claude Opus 4.6, with OpenAI’s GPT-5.1 run on a subset for comparison. So the model was not guessing at strangers. It was given each person’s demographics and a long list of their previous answers, and was asked what that person would say next. A persona built from a sketch of a customer segment starts with far less.

How far off, and in which direction#

“Across nearly 300 individual survey questions, the estimates produced using our AI respondents differed from their human counterparts by an average of 12 percentage points.” On about 28 per cent of questions the gap passed 15 points. The examples show the shape of the error. Real approval of the President stood at 34 per cent and the twins put it at 46. A quarter of real respondents had heard a lot about data centres and 3 per cent of the twins had. On a question about what the First Amendment guarantees, the figures were 52 per cent for people and 98 for the twins. On knowledge the synthetic public looks better informed than the real one. It is also tidier throughout: “Nearly half the questions we asked had at least one answer choice that was not selected by a single AI-generated respondent.” It also exaggerates groups. “In many cases, our AI poll took beliefs or attitudes that are reasonably common among a particular subgroup and portrayed them as nearly ubiquitous.” Average error was highest for Republicans, at 16.1 points, and for Black adults, at 15.1; the report’s own summary is that synthetic surveys are “especially error-prone for some subgroups”.

Two models, two publics#

The comparison between laboratories is the part a buyer should read twice. On the same questions with the same instructions, “the GPT estimates described an American public that has more extreme opinions on a variety of topics than they actually do, while the Opus estimates described a public that is more middle-of-the-road than in reality. A reader of either result would be misled, but in different directions.” Neither model is singled out as worse, and the estate does not rank them. The point is that the answer depended on which product was asked, and a reader with only one of the two would have had no way to see it. A survey of real people can be wrong, and its errors can be estimated from the sample. A synthetic survey’s error is a property of a model that may be replaced by another before the next wave.

Where Pew draws its own line#

Two days earlier Claudia Deane set out how the Center does and does not use AI, and the document is a usable model of a policy written by task. “Humans decide what we study.” “We don’t use AI to create or model synthetic public opinion. Our survey results are based on the views that real people report to us.” “Humans oversee every aspect of our work, and humans hold the ultimate responsibility for its quality and accuracy.” On the other side of the line, engineers use code assistants, and AI helps sort open-ended answers into categories and takes a first pass at copy editing. A different use of AI in research keeps the people in. Anthropic’s study, opened on 29 September, uses an AI interviewer to question real users, and states its own limit: “Participants in this study will all be people who use Claude, which is not a representative sample of the public.” A machine asking people questions and a machine answering for them are different decisions, and only the second removes the people.

Why it matters outside polling#

The same technique is offered to companies under names such as synthetic customers and AI personas. Inside an organisation it is a tempting substitute for asking staff. Pew’s three errors map onto what such research is for. A survey is commissioned to find the minority view, and the twins erased rare answers. It is commissioned to learn what people do not know, and the twins knew nearly everything. It is commissioned to hear from particular groups, and for particular groups the error was largest. There is a second cost. A team that stops talking to customers because a model can stand in for them loses the practice of listening, and the estate’s page on the absent user describes what a decision looks like when the people affected have no voice in the room. A synthetic respondent is an absent user with a confident spokesman. The wider pattern, a model pulling answers towards the middle, is examined on whether AI makes everyone think alike.

The decision an organisation controls#

Rules Before Tools turns Pew’s line into four answers. Which decisions may rest on a simulated respondent: drafting a questionnaire and testing its wording, yes; a figure in a board paper about what customers or employees think, no. Who can stop one: whoever signs the paper, helped by a labelling rule that every number about people says whether people were asked. What must people remain able to do: run a real survey and hold a real conversation with a customer, skills that go when a cheaper stand-in is always to hand. And how anyone would know it had gone wrong: put a sample of the same questions to real people at intervals and publish the gap, which is the test Pew ran. The measurable risk sits in the handover. A model’s estimate of opinion, once it is in a chart, looks the same as a finding, and what happens to human judgement afterwards depends on whether anyone can still tell the two apart.

What this does not show#

This is one organisation’s test of two models on American political and social questions in early 2026, and Pew’s own conclusion carries the words “at this time”. The report does not state how many panel members were twinned. It is not peer reviewed. It does not test commercial tools trained on a company’s own customer data, which may do better or worse, and it says nothing about questions of purchase behaviour. It does not test AI as interviewer. The largest gaps quoted here are on knowledge questions, and an average hides items where the twins were close. Nothing here shows that decisions made on synthetic research turned out worse, because nobody has measured that.

Essay · SS-2026-378 · 1 institutional survey

Cite this page

Hirji, R. (2026). Can AI stand in for real people in a survey?. The SuperSkills evidence base, SS-2026-378. https://thesuperskills.com/research/can-ai-stand-in-for-real-people-in-a-survey. Last reviewed 2 October 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

Can AI stand in for real people in a survey?

Not yet. Pew Research Center reported on 30 September 2026 that AI respondents modelled on its own panel members differed from the real answers by an average of 12 percentage points across nearly 300 questions, and by more than 15 points on about 28 per cent of them. Pew's conclusion is that AI polling is not a replacement for rigorously surveying real humans.

How accurate are synthetic survey respondents?

In Pew's test, off by 12 points on average, with the error larger for some groups: 16.1 points for Republicans and 15.1 for Black adults. On nearly half the questions at least one answer was chosen by no AI respondent, and on knowledge questions the AI respondents scored far higher than people. Two models erred in opposite directions, one describing a more extreme public and one a more moderate one.

Should a company use AI personas instead of asking customers?

Not as the finding. Pew's errors fall where research earns its cost: rare views disappear, people appear better informed than they are, and some groups are described worst. A simulated respondent can help draft and test a questionnaire. A figure about what customers or employees think should say whether people were asked.

Does Pew Research Center use AI in its polls?

Not to generate opinion. Its statement of 28 September 2026 says humans decide what it studies and write and review its reports, and that it does not use AI to create or model synthetic public opinion. It uses AI for writing code, for sorting open-ended answers into categories and in the first stage of copy editing.

In this hub

Judgement, oversight and accountability

Who decides, who checks, and who is answerable when the machine was involved.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

One rule for every figure about people that reaches a decision-maker, whether real people were asked, which uses of simulated respondents are allowed, who can stop a synthetic number entering a board paper, and a periodic test of the simulation against a real sample, is the engagement. AI advisory for CEOs and boards.

This is the question underneath the rest of them, and the one a leadership team is least likely to have named out loud. There is Speaker on AI and human judgement, and the full range of topics and audiences.

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.