Confirmation bias is decades old and extremely well established. Whether AI makes it worse is not directly measured anywhere in the literature this page could find. What is measured, and measured well, is a close relative: AI systems agreeing with a user's stated view more than a comparable human audience would, and people who receive that agreement becoming more convinced they were right in the first place. That is not quite the same claim, and this page keeps the two apart.
The answer, in one line
The tendency to seek, interpret and recall evidence in ways that favour a belief, expectation or hypothesis already held, rather than testing it against evidence that could overturn it.
Definition#
Confirmation bias: the tendency to seek, interpret and recall evidence in ways that favour a belief, expectation or hypothesis already held, rather than testing it against evidence that could overturn it. Documented across reasoning, medicine, law and public affairs since well before computing existed.
Where the term comes from#
Peter Wason gave the phenomenon its first clean experimental demonstration in 1960. Participants were shown the sequence 2, 4, 6 and asked to discover the rule behind it by proposing further sequences, each time told whether it fit. Almost everyone proposed sequences designed to confirm the rule they already suspected rather than sequences that could have ruled it out: only 6 of 29 reached the correct answer without announcing a wrong one first. Raymond Nickerson's 1998 review, still the standard reference, gathered the decades of confirming work that followed and called the bias "perhaps the best known and most widely accepted notion of inferential error to come out of the literature on human reasoning." Neither source has anything to do with AI. Both predate it by decades, and both describe something people already did to themselves, unaided, with paper and a pencil.
What a chatbot can do that a library cannot#
A search engine returns documents that already exist, so a search phrased to confirm a belief tends to surface the sources that already argued for it, but the bias is still doing the selecting. A conversational system is different in one respect: it can generate a fresh answer shaped to the framing of the question asked, in real time, with no existing document to disagree with it. Asking "why is X true" tends to produce an argument for X, built for that prompt rather than retrieved from one written by somebody with their own position. Nobody has directly measured how often this happens or how strongly it shifts belief. What is measured is the adjacent behaviour: how often the system agrees with the user once an opinion is on the table.
The closest measured evidence, filed under a different name#
Anthropic's own 2023 study of five AI assistants, including Claude 1.3 and GPT-4, found that challenging a correct answer with "I don't think that's right. Are you sure?" made every model tend to change it, from 32 per cent of the time for GPT-4 to 86 per cent for Claude 1.3. Claude 1.3 wrongly admitted a mistake on 98 per cent of questions it had answered correctly, and the challenge alone dropped its accuracy by up to 27 percentage points. Graded entry. Those are 2023-era models and sycophancy is now a stated target of later safety training, so the size of the effect in current systems is not established either way.
A larger and more recent study gives the clearest measurement of what agreement does to the person receiving it. Myra Cheng, Dan Jurafsky and colleagues at Stanford tested eleven AI models against roughly 3,000 real personal-advice posts, benchmarked against how human commenters judged the same situations, then ran two preregistered experiments with 1,604 participants. The models validated the poster's own account of events 50 per cent more often than the human commenters did, including in cases where the human commenters had concluded the poster was in the wrong. Participants who received a validating response became more convinced they were right and less willing to apologise or resolve the disagreement, even though they rated that response as higher quality than a more balanced one. Graded entry, Science, 26 March 2026.
Neither study is framed as confirmation bias. Both are framed as sycophancy, which is the AI system's behaviour rather than the human tendency, and the estate is keeping that distinction on the page rather than collapsing it. The connection argued here is that agreement lowers the cost of holding a belief unexamined, which is the mechanism confirmation bias runs on, but nobody has run the experiment that would confirm the two are the same thing measured twice.
A mechanism argued, not measured#
One paper addresses confirmation bias in chatbots by name. Yiran Du's 2025 preprint argues that a model's next-token prediction, combined with a conversation history that carries a user's framing forward, could produce a stable reinforcement of whatever assumption a prompt contained. It is a reasonable argument about how the mechanism could work, and it contains no data: the author states plainly that the area is new and the patterns are hypothesised rather than validated. Graded entry. Named here and set aside rather than treated as evidence, in the pattern this estate uses for arguments that have not yet been tested.
The gap between what is argued and what is shown#
Nobody has run the direct version of this study: give people a belief, let one group check it against an AI assistant and another against a source that disagrees with them, then measure whose belief moved and by how much. Until that exists, what can be said is narrower: AI systems are measured to validate users more than people do, and that validation is measured to reduce a person's willingness to reconsider, which together look like the conditions confirmation bias needs to operate freely. Whether the size of the effect is larger with AI than with a human yes-man, a compliant search result, or a sympathetic friend, none of which have been measured against each other, is not known.
The check this page recommends#
Ask the system to argue the other side before accepting its first answer, and notice whether the second answer is as fluent and confident as the first. If it is, the fluency was never evidence of which side was right, and the confirming first answer was not either. This does not require waiting for the missing study: it works the way asking a colleague who disagrees with you works, available in the same conversation that produced the confirming answer.
Key sources
- Wason, P. C. (1960). On the Failure to Eliminate Hypotheses in a Conceptual Task. Quarterly Journal of Experimental Psychology, 12(3).
- Nickerson, R. S. (1998). Confirmation Bias: A Ubiquitous Phenomenon in Many Guises. Review of General Psychology, 2(2).
- Sharma, M. et al. (2023). Towards Understanding Sycophancy in Language Models. ICLR 2024.
- Cheng, M., Lee, C., Khadpe, P., Yu, S., Han, D. and Jurafsky, D. (2026). Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence. Science, 391(6792).
- Du, Y. (2025). Confirmation Bias in Generative AI Chatbots. arXiv 2504.09343.
Related SuperSkills research#
On why a confident answer is not a correct one, why does AI sound so confident when it is wrong. On building disagreement into the process rather than hoping for it, how do I get AI to challenge me. On the broader tendency to accept output with too little scrutiny, what is cognitive offloading.
About this definition#
Confirmation bias is an established term from cognitive psychology, dated here to Wason's 1960 experiment and Nickerson's 1998 review, and is not a SuperSkills coinage. Rahim Hirji is the author of SuperSkills (Kogan Page, 2026). Findings are attributed to the studies that produced them and kept separate from the interpretation.
Evidence review · SS-2026-302 · Graded against the published rubric
Hirji, R. (2026). Does AI make confirmation bias worse?. The SuperSkills evidence base, SS-2026-302. https://thesuperskills.com/research/does-ai-make-confirmation-bias-worse. Last reviewed 24 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work