← Research
Research

Should I let AI summarise everything I read?

The cost does not show up in what you feel you know. It shows up in what you can then make out of it, and in how much your version resembles everyone else's.

Last reviewed: 28 August 2026

What Melumad and Yun measured when they held the facts constant, why searching produces an illusion of knowledge even when it returns nothing, and the three kinds of reading where a summary is the wrong instrument.

Not everything, and the reason is now measured rather than asserted. People who learn a topic from an AI summary instead of from links spend less time on it, report shallower knowledge, and then produce work that is shorter, carries fewer specific facts, and looks far more like everyone else's. The most useful experiment held the facts identical across both conditions, so what changed was the format, not the information. The summary is a good instrument for deciding whether to read something. It is a poor one for reading it.

Seven experiments, and the one that isolates the format

Melumad and Yun published seven experiments in PNAS Nexus in October 2025, four in the paper and three in the supplement. Participants learned about a practical topic, planting a vegetable garden, leading a healthier lifestyle, or what to do after a financial scam, using either a large language model or web search links, and then wrote advice for a friend.

Experiment 1, with 1,104 participants using the real ChatGPT and the real Google, found the pattern. The AI group spent 585.41 seconds against 742.81 for the search group. They rated themselves lower on having learned new things (3.43 against 3.86 on a five-point scale) and on ownership of what they learned (3.36 against 3.55). Their advice was shorter (84.58 words against 94.64) and carried fewer references to specific entities (0.464 against 0.718). And the advice was much more alike: mean pairwise cosine similarity of 0.159 against 0.057.

The obvious objection is that ChatGPT and Google surface different information. Experiment 2 removed it. The authors generated a 291-word synthesis of seven suggestions, then had the model rewrite the same facts as six articles in different publication styles, presented as six links. Same facts, two formats, 1,979 participants.

The effects held, and several got larger.

Read the third item alongside the rest. People given the summary rated it exactly as comprehensive as the links, and then wrote something thinner and more generic. The instrument did not feel worse while it was being used.

Experiment 3 closed the other escape route. It held the search engine constant, comparing standard Google against Google with AI Overviews, with 250 lab participants. Engagement time fell from 67.37 to 55.50 seconds and rated comprehensiveness from 4.18 to 3.32. So the effect is not about the novelty of a chatbot interface.

The number the authors could not make people use

One robustness check in the supplement deserves more attention than it gets. The researchers tested a condition where the AI summary included real-time web links, the obvious fix. The effects persisted, and the authors explain why: only 26 per cent of those participants clicked any of the links.

Offering the source is not the same as anyone opening it. The design most products have converged on, a summary with citations underneath, was tested and did not solve the problem, because a summary that answers the question removes the reason to go further.

Why it feels like understanding

This has been measured for two decades under a different name. Fisher, Goddu and Keil ran nine experiments with 1,708 participants. Half were asked to look up explanations online, half were not, and everyone then rated how well they could explain questions in six domains unrelated to what they had searched for.

The searchers rated themselves higher, with Cohen's d between 0.35 and 0.63 across the studies. Two of the follow-ups are the memorable ones. In Experiment 4b, the illusion appeared just as strongly when the search engine returned no answer to the question asked (4.11 against 4.00 for people who found an answer, a difference of nothing, both well above the no-search baseline of 3.05). In Experiment 4c it survived when the search returned no results at all.

And it is specific rather than general. Experiment 3 used autobiographical questions, where the internet is no help, and the effect vanished. So this is not vague overconfidence. It is a targeted miscalibration in exactly the domains where information is available on demand.

The authors are careful about what they did not test: every measure is a self-rating, and no experiment assessed whether people's actual explanatory ability changed. The finding is about the gap between felt and held knowledge, and that gap is the thing a summary widens.

The mechanism the older literature already named

Sparrow, Liu and Wegner showed in 2011 that when people expect information to remain available, they encode where to find it rather than the thing itself. Risko and Gilbert's review establishes that the decision to offload is metacognitive, driven by how hard a task feels, and frequently mistaken.

Put the three together and the summary problem has a shape. The summary lowers the felt difficulty of the task, which is the trigger for offloading. It leaves the impression of comprehensiveness, which removes the signal that anything is missing. And it withholds the specific material that would have been encoded, which shows up later as thinner output. None of that is available to introspection while it is happening, and that is the reason a self-check on this fails.

The learning literature has a name for what has been removed. Effortful retrieval and self-organised search are desirable difficulties: they slow performance during learning and improve what survives.

Three kinds of reading, and only one of them summarises well

The third category is smaller than people think and gets treated as the first. The practical failure mode is not summarising too much overall. It is summarising the wrong ten per cent.

What this evidence does not establish

Six habits that follow from the numbers

Key research and primary sources

Related SuperSkills research

On the mechanism, cognitive offloading, the Google effect and desirable difficulty. On the effect on thinking, AI and critical thinking and does AI make everyone think alike. On memory specifically, do I still need to remember things. On dependency, am I becoming dependent on AI and using AI without dependency.

About this research

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026) and founder of The SuperSkills Intelligence Company. Every figure from Melumad and Yun and from Fisher, Goddu and Keil was read at the primary source. Experiment 4 of the Melumad paper, covering recipients' willingness to adopt the advice, could not be retrieved in full at the publisher, so its direction is reported from the abstract and no figures from it appear here. Reviewed quarterly.

Cite this

Hirji, R. (2026). Should I let AI summarise everything I read? The SuperSkills Intelligence Company. Last reviewed 28 August 2026. thesuperskills.com/research/should-i-let-ai-summarise-everything-i-read

In this hub

AI and Human Judgement

Does AI weaken judgement? The evidence, and what to do about it.

The work

Where the writing comes from.

These essays draw on research across more than 200 organisations in 30 countries. See the wider body of work, or bring it into your organisation.

All research →