Not everything, and the reason is now measured rather than asserted. People who learn a topic from an AI summary instead of from links spend less time on it, report shallower knowledge, and then produce work that is shorter, carries fewer specific facts, and looks far more like everyone else's. The most useful experiment held the facts identical across both conditions, so what changed was the format, not the information. The summary is a good instrument for deciding whether to read something. It is a poor one for reading it.
Seven experiments, and the one that isolates the format
Melumad and Yun published seven experiments in PNAS Nexus in October 2025, four in the paper and three in the supplement. Participants learned about a practical topic, planting a vegetable garden, leading a healthier lifestyle, or what to do after a financial scam, using either a large language model or web search links, and then wrote advice for a friend.
Experiment 1, with 1,104 participants using the real ChatGPT and the real Google, found the pattern. The AI group spent 585.41 seconds against 742.81 for the search group. They rated themselves lower on having learned new things (3.43 against 3.86 on a five-point scale) and on ownership of what they learned (3.36 against 3.55). Their advice was shorter (84.58 words against 94.64) and carried fewer references to specific entities (0.464 against 0.718). And the advice was much more alike: mean pairwise cosine similarity of 0.159 against 0.057.
The obvious objection is that ChatGPT and Google surface different information. Experiment 2 removed it. The authors generated a 291-word synthesis of seven suggestions, then had the model rewrite the same facts as six articles in different publication styles, presented as six links. Same facts, two formats, 1,979 participants.
The effects held, and several got larger.
- Time engaging with results: 83.65 seconds with the summary against 124.32 seconds with links.
- Learned new things: 3.71 against 3.96. Ownership of the knowledge: 3.41 against 3.66.
- Comprehensiveness: no difference at all (4.30 against 4.25, p = 0.264). The summary felt just as complete.
- Thought and effort put into the advice: 3.85 against 4.11.
- Advice length: 64.49 words against 74.22. References to specific facts: 4.00 against 4.61.
- Similarity to other participants' advice: 0.224 against 0.072, a threefold increase in how alike the outputs were.
Read the third item alongside the rest. People given the summary rated it exactly as comprehensive as the links, and then wrote something thinner and more generic. The instrument did not feel worse while it was being used.
Experiment 3 closed the other escape route. It held the search engine constant, comparing standard Google against Google with AI Overviews, with 250 lab participants. Engagement time fell from 67.37 to 55.50 seconds and rated comprehensiveness from 4.18 to 3.32. So the effect is not about the novelty of a chatbot interface.
The number the authors could not make people use
One robustness check in the supplement deserves more attention than it gets. The researchers tested a condition where the AI summary included real-time web links, the obvious fix. The effects persisted, and the authors explain why: only 26 per cent of those participants clicked any of the links.
Offering the source is not the same as anyone opening it. The design most products have converged on, a summary with citations underneath, was tested and did not solve the problem, because a summary that answers the question removes the reason to go further.
Why it feels like understanding
This has been measured for two decades under a different name. Fisher, Goddu and Keil ran nine experiments with 1,708 participants. Half were asked to look up explanations online, half were not, and everyone then rated how well they could explain questions in six domains unrelated to what they had searched for.
The searchers rated themselves higher, with Cohen's d between 0.35 and 0.63 across the studies. Two of the follow-ups are the memorable ones. In Experiment 4b, the illusion appeared just as strongly when the search engine returned no answer to the question asked (4.11 against 4.00 for people who found an answer, a difference of nothing, both well above the no-search baseline of 3.05). In Experiment 4c it survived when the search returned no results at all.
And it is specific rather than general. Experiment 3 used autobiographical questions, where the internet is no help, and the effect vanished. So this is not vague overconfidence. It is a targeted miscalibration in exactly the domains where information is available on demand.
The authors are careful about what they did not test: every measure is a self-rating, and no experiment assessed whether people's actual explanatory ability changed. The finding is about the gap between felt and held knowledge, and that gap is the thing a summary widens.
The mechanism the older literature already named
Sparrow, Liu and Wegner showed in 2011 that when people expect information to remain available, they encode where to find it rather than the thing itself. Risko and Gilbert's review establishes that the decision to offload is metacognitive, driven by how hard a task feels, and frequently mistaken.
Put the three together and the summary problem has a shape. The summary lowers the felt difficulty of the task, which is the trigger for offloading. It leaves the impression of comprehensiveness, which removes the signal that anything is missing. And it withholds the specific material that would have been encoded, which shows up later as thinner output. None of that is available to introspection while it is happening, and that is the reason a self-check on this fails.
The learning literature has a name for what has been removed. Effortful retrieval and self-organised search are desirable difficulties: they slow performance during learning and improve what survives.
Three kinds of reading, and only one of them summarises well
- Reading to decide whether to read. Triage. A summary is the correct instrument and there is no argument against it. Most inbox and feed reading is this.
- Reading to use once. A council notice, a product manual, a policy you will follow and forget. Summarise it. Nothing is being built.
- Reading that feeds your own judgement. Anything you will argue from, decide on, be accountable for, or be expected to notice an error in. Here the compression removes the things the evidence says get removed: the specific facts, the distinctive framings, the sense of where the argument is weak. On the similarity measure, it also removes what made your version different from everyone else's.
The third category is smaller than people think and gets treated as the first. The practical failure mode is not summarising too much overall. It is summarising the wrong ten per cent.
What this evidence does not establish
- Depth of learning was self-reported in every experiment. There was no objective recall or comprehension test. What is measured objectively is time, word count, factual density and similarity, all of which are consistent with the self-reports but are not the same claim.
- Time on task is a proxy for effort, and the authors label it as one.
- The recipient-side results are not quoted here. The paper reports that recipients were less likely to adopt advice written after AI use. This research could not retrieve Experiment 4 in full at the publisher and therefore gives the direction without the figures, rather than repeating numbers it has not read.
- The paper states its own total sample twice, as 10,462 in the abstract and 10,426 in the introduction. Both are printed. The per-experiment figures quoted above are the ones read directly.
- The tasks were practical how-to topics with short horizons. Whether the same holds for reading a contract, a novel or a scientific paper is untested.
- Nobody has tested a countermeasure. Reading the source afterwards, asking for the disagreements rather than the summary, or having the model quiz you have all been proposed and none measured.
Six habits that follow from the numbers
- Decide before you summarise which of the three kinds of reading this is. That single question does most of the work.
- Do not trust the feeling of comprehensiveness. It was identical across conditions in the experiment where the output was measurably thinner.
- Assume you will not click the links. Seventy-four per cent of people did not, when the links were right there.
- Ask for the disagreements, not the summary. A summary of what a text argues removes the friction; a list of where sources conflict restores some of it. Untested, and stated as untested.
- Notice the similarity effect. If your view of a subject came from a summary, so did everyone else's, and it came from the same one.
- Read the primary source for anything you will be held to. The specific facts are the first thing compression removes and the first thing an expert asks for.
Key research and primary sources
- Melumad, S. and Yun, J. H. (2025). Experimental evidence of the effects of large language models versus web search on depth of learning. PNAS Nexus, 4(10), pgaf316.
- Fisher, M., Goddu, M. K. and Keil, F. C. (2015). Searching for explanations: How the Internet inflates estimates of internal knowledge. Journal of Experimental Psychology: General, 144(3).
- Sparrow, B., Liu, J. and Wegner, D. M. (2011). Google Effects on Memory. Science, 333(6043).
- Risko, E. F. and Gilbert, S. J. (2016). Cognitive Offloading. Trends in Cognitive Sciences, 20(9).
- Lee, H.-P. et al. (2025). The Impact of Generative AI on Critical Thinking. Microsoft Research and Carnegie Mellon, CHI 2025.
Related SuperSkills research
On the mechanism, cognitive offloading, the Google effect and desirable difficulty. On the effect on thinking, AI and critical thinking and does AI make everyone think alike. On memory specifically, do I still need to remember things. On dependency, am I becoming dependent on AI and using AI without dependency.
About this research
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026) and founder of The SuperSkills Intelligence Company. Every figure from Melumad and Yun and from Fisher, Goddu and Keil was read at the primary source. Experiment 4 of the Melumad paper, covering recipients' willingness to adopt the advice, could not be retrieved in full at the publisher, so its direction is reported from the abstract and no figures from it appear here. Reviewed quarterly.
Cite this
Hirji, R. (2026). Should I let AI summarise everything I read? The SuperSkills Intelligence Company. Last reviewed 28 August 2026. thesuperskills.com/research/should-i-let-ai-summarise-everything-i-read