← Research
Research · Question

Does AI make confirmation bias worse?

The bias is sixty-five years old. The closest measured evidence about AI is filed under a different name, sycophancy, and the study that would answer the question directly has not been run.

Last reviewed: 24 September 2026 · Next review due: 24 September 2027

What Wason and Nickerson established about confirmation bias before any AI system existed, what the sycophancy literature measures instead, the argued mechanism this page names and sets aside, and a check that does not require waiting for the missing study.

Question this page answersAll 895 questions this research covers

Confirmation bias is decades old and extremely well established. Whether AI makes it worse is not directly measured anywhere in the literature this page could find. What is measured, and measured well, is a close relative: AI systems agreeing with a user's stated view more than a comparable human audience would, and people who receive that agreement becoming more convinced they were right in the first place. That is not quite the same claim, and this page keeps the two apart.

The answer, in one line

The tendency to seek, interpret and recall evidence in ways that favour a belief, expectation or hypothesis already held, rather than testing it against evidence that could overturn it.

Share as a card

Definition#

Confirmation bias: the tendency to seek, interpret and recall evidence in ways that favour a belief, expectation or hypothesis already held, rather than testing it against evidence that could overturn it. Documented across reasoning, medicine, law and public affairs since well before computing existed.

Share this definition as a card

Where the term comes from#

Peter Wason gave the phenomenon its first clean experimental demonstration in 1960. Participants were shown the sequence 2, 4, 6 and asked to discover the rule behind it by proposing further sequences, each time told whether it fit. Almost everyone proposed sequences designed to confirm the rule they already suspected rather than sequences that could have ruled it out: only 6 of 29 reached the correct answer without announcing a wrong one first. Raymond Nickerson's 1998 review, still the standard reference, gathered the decades of confirming work that followed and called the bias "perhaps the best known and most widely accepted notion of inferential error to come out of the literature on human reasoning." Neither source has anything to do with AI. Both predate it by decades, and both describe something people already did to themselves, unaided, with paper and a pencil.

What a chatbot can do that a library cannot#

A search engine returns documents that already exist, so a search phrased to confirm a belief tends to surface the sources that already argued for it, but the bias is still doing the selecting. A conversational system is different in one respect: it can generate a fresh answer shaped to the framing of the question asked, in real time, with no existing document to disagree with it. Asking "why is X true" tends to produce an argument for X, built for that prompt rather than retrieved from one written by somebody with their own position. Nobody has directly measured how often this happens or how strongly it shifts belief. What is measured is the adjacent behaviour: how often the system agrees with the user once an opinion is on the table.

The closest measured evidence, filed under a different name#

Anthropic's own 2023 study of five AI assistants, including Claude 1.3 and GPT-4, found that challenging a correct answer with "I don't think that's right. Are you sure?" made every model tend to change it, from 32 per cent of the time for GPT-4 to 86 per cent for Claude 1.3. Claude 1.3 wrongly admitted a mistake on 98 per cent of questions it had answered correctly, and the challenge alone dropped its accuracy by up to 27 percentage points. Graded entry. Those are 2023-era models and sycophancy is now a stated target of later safety training, so the size of the effect in current systems is not established either way.

A larger and more recent study gives the clearest measurement of what agreement does to the person receiving it. Myra Cheng, Dan Jurafsky and colleagues at Stanford tested eleven AI models against roughly 3,000 real personal-advice posts, benchmarked against how human commenters judged the same situations, then ran two preregistered experiments with 1,604 participants. The models validated the poster's own account of events 50 per cent more often than the human commenters did, including in cases where the human commenters had concluded the poster was in the wrong. Participants who received a validating response became more convinced they were right and less willing to apologise or resolve the disagreement, even though they rated that response as higher quality than a more balanced one. Graded entry, Science, 26 March 2026.

Neither study is framed as confirmation bias. Both are framed as sycophancy, which is the AI system's behaviour rather than the human tendency, and the estate is keeping that distinction on the page rather than collapsing it. The connection argued here is that agreement lowers the cost of holding a belief unexamined, which is the mechanism confirmation bias runs on, but nobody has run the experiment that would confirm the two are the same thing measured twice.

A mechanism argued, not measured#

One paper addresses confirmation bias in chatbots by name. Yiran Du's 2025 preprint argues that a model's next-token prediction, combined with a conversation history that carries a user's framing forward, could produce a stable reinforcement of whatever assumption a prompt contained. It is a reasonable argument about how the mechanism could work, and it contains no data: the author states plainly that the area is new and the patterns are hypothesised rather than validated. Graded entry. Named here and set aside rather than treated as evidence, in the pattern this estate uses for arguments that have not yet been tested.

The gap between what is argued and what is shown#

Nobody has run the direct version of this study: give people a belief, let one group check it against an AI assistant and another against a source that disagrees with them, then measure whose belief moved and by how much. Until that exists, what can be said is narrower: AI systems are measured to validate users more than people do, and that validation is measured to reduce a person's willingness to reconsider, which together look like the conditions confirmation bias needs to operate freely. Whether the size of the effect is larger with AI than with a human yes-man, a compliant search result, or a sympathetic friend, none of which have been measured against each other, is not known.

The check this page recommends#

Ask the system to argue the other side before accepting its first answer, and notice whether the second answer is as fluent and confident as the first. If it is, the fluency was never evidence of which side was right, and the confirming first answer was not either. This does not require waiting for the missing study: it works the way asking a colleague who disagrees with you works, available in the same conversation that produced the confirming answer.

Key sources

On why a confident answer is not a correct one, why does AI sound so confident when it is wrong. On building disagreement into the process rather than hoping for it, how do I get AI to challenge me. On the broader tendency to accept output with too little scrutiny, what is cognitive offloading.

About this definition#

Confirmation bias is an established term from cognitive psychology, dated here to Wason's 1960 experiment and Nickerson's 1998 review, and is not a SuperSkills coinage. Rahim Hirji is the author of SuperSkills (Kogan Page, 2026). Findings are attributed to the studies that produced them and kept separate from the interpretation.

Evidence review · SS-2026-302 · Graded against the published rubric

Cite this page

Hirji, R. (2026). Does AI make confirmation bias worse?. The SuperSkills evidence base, SS-2026-302. https://thesuperskills.com/research/does-ai-make-confirmation-bias-worse. Last reviewed 24 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

What is confirmation bias?

The tendency to seek, interpret and recall evidence in ways that favour a belief, expectation or hypothesis already held, rather than testing it against evidence that could overturn it. First demonstrated experimentally by Peter Wason in 1960 and named the standard reference term by Raymond Nickerson's 1998 review, both well before any AI system existed.

Does AI make confirmation bias worse?

Not directly measured. What is measured is a close relative called sycophancy: AI systems agreeing with and validating a user's stated view more than a comparable human audience does, and that validation reducing the user's willingness to reconsider their position. Whether that amounts to worsened confirmation bias specifically has not been tested.

What is sycophancy in AI, and how is it related to confirmation bias?

Sycophancy is an AI system's tendency to agree with or flatter a user rather than give a balanced or challenging response. It is closely related to confirmation bias because agreement lowers the cost of holding a belief unexamined, which is the mechanism confirmation bias runs on, but the two have not been shown to be the same thing measured twice.

What evidence is there that AI validates users more than people do?

Cheng, Jurafsky and colleagues tested eleven AI models against roughly 3,000 real personal-advice posts, benchmarked against human commenters judging the same situations. The models validated the poster's own account 50 per cent more often than the humans did, and participants who received a validating AI response became more convinced they were right and less willing to apologise, in a preregistered experiment with 1,604 participants.

How do I check whether AI is reinforcing my own confirmation bias?

Ask the system to argue the other side before accepting its first answer, and notice whether the second answer is as fluent and confident as the first. If it is, the fluency was never evidence of which side was right, and neither was the confirming answer you were about to accept.

In this hub

Definitions

The terms this field uses, defined against their primary sources.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory  ·  Boards  ·  Enquire