← Research
Research · Question

Is AI making us stupid?

Does AI make you lazy, worse at thinking, or bad for your brain? Three different questions, and only two of them have been measured.

Last reviewed: 23 September 2026 · Next review due: 23 September 2027

The question collapses three separable claims. AI changes what people bother to remember, which is measured. AI leaves people worse at a task once it is withdrawn, which is measured in schools, laboratories and an endoscopy suite. AI lowers general cognitive ability, which nothing measures at all. The Norwegian conscription record that does show a population decline puts its turning point at the 1975 birth cohort.

Questions this page answersAll 887 questions this research covers

Nobody has measured it. There is no study anywhere that takes a person's general intelligence, gives them an AI assistant for a year, and measures it again. What has been measured, repeatedly and by several methods, is something narrower and more useful: how well people do a task once the assistance is taken away. That answer is consistent, uncomfortable, and much more specific than the question as people ask it.

The answer, in one line

Nobody has measured it. No study takes a person's general cognitive ability, gives them an AI assistant, and measures it again.

Share as a card

Definition#

Is AI making us stupid: a question that collapses three separable claims. First, that AI changes what people bother to learn or remember, which is measured and true. Second, that AI leaves people worse at a task once it is withdrawn, which is measured in several settings and true there. Third, that AI lowers general cognitive ability, which nothing measures at all. Only the third claim is what the word stupid usually means.

Share this definition as a card

Three floors, and only two have anything under them#

The first floor is the oldest and the least contested. In 2011 Sparrow, Liu and Wegner ran four laboratory experiments on what people encode when they expect information to stay available. People who believed a fact would remain on the computer remembered where to find it rather than the fact. That is a trade in what gets stored, and the authors were careful to say it does not show total memory capability falling. Search engines did this before chat assistants existed.

The second floor is the one the last two years have filled in. Take the tool away and test the person, and a penalty appears. That describes a practice rather than a person, and it carries most of the evidence behind the popular question.

The third floor is empty. General cognitive ability, the thing an IQ test tries to estimate, has not been measured as a function of AI exposure by anybody. No cohort, no trial, no panel. A page claiming otherwise is reading one of the first two floors and labelling it as the third.

The population score everyone reaches for started falling in 1975#

Measured cognitive scores have fallen across a whole population in one well-documented case, and the date the fall began settles a lot. Bernt Bratsberg and Ole Rogeberg, working from Norwegian military conscription records, analysed birth cohorts from 1962 to 1991. Their sample was 817,611 men present in Norway on their eighteenth birthday, of whom 736,808 had a valid ability score. Mean IQ rose from 99.5 for the 1962 cohort to 102.3 for the 1975 cohort, then fell back to 99.4 by 1989.

What makes the paper matter is the design. They recovered the rise, the turning point and the fall from variation within families, comparing brothers. That rules out explanations depending on who was having children, which is where the debate had been stuck. Their selection-corrected estimate for the decline period is 0.33 IQ points per year within families, against 0.34 across families. The Flynn effect and its reversal are both environmental, and the environment doing it is shared by siblings.

The 1991 cohort sat that test around 2009. Every person in the falling half of that curve was assessed before generative AI existed. Whatever has been moving those scores for fifty years, it is not ChatGPT. The authors are equally clear about the limit of their own result: they cannot identify which environmental factor is responsible, and they list changing media exposure among the hypotheses their design leaves standing. So the paper does not exonerate screens. It dates the clock.

The experiments that withdraw the tool and then test you#

Four results, four populations, one shape.

Bastani and colleagues ran a field experiment with nearly 1,000 high-school mathematics students in three arms: unrestricted GPT-4, a hints-only tutor, and a control. With the tool present, grades rose 48 per cent under unrestricted access and 127 per cent with the tutor. With access removed, the unrestricted group scored 17 per cent lower than students who had never had it. The guardrailed tutor largely removed the harm. The interface decided the outcome; the presence of AI did not.

Liu and colleagues produced the effect causally and fast. Across randomised trials with 1,222 participants on mathematical reasoning and reading comprehension, assistance was available during practice and then withdrawn. People performed worse unassisted and were more likely to give up, after roughly ten minutes of exposure. Their reading moves the mechanism from knowledge to persistence, which travels further than a subject-specific result.

Shen and Tamkin tested 52 experienced Python developers learning an unfamiliar library, half with an AI sidebar, then quizzed them without it. The assisted arm averaged 50 per cent against 67 per cent, a pre-registered result at p = 0.010. The control arm hit more errors, with a median of 3.0 against 1.0, and got better at resolving them. Most of the gap sat in the debugging questions.

And outside a laboratory entirely, Budzyn and colleagues looked at 1,443 colonoscopies performed without AI assistance by 19 Polish endoscopists averaging 27.6 years of experience, before and after AI was introduced at their centres. Unassisted adenoma detection fell from 28.4 per cent to 22.4 per cent, a drop of 6.0 percentage points. It is observational rather than randomised, and the authors say so. It is also the only measurement here on people with decades of expertise doing paid clinical work.

None of these measures intelligence. All of them measure what a person can do alone after a period of not doing it alone.

The three studies that get quoted are the three weakest ones here#

Michael Gerlich surveyed and interviewed 666 people and found a negative correlation between frequent AI use and critical-thinking scores, mediated by cognitive offloading and strongest among the youngest. It is a correlation with a plausible mechanism and no causal claim, and it carries a published correction issued on 10 September 2025 which anybody citing it should read alongside.

Lee and colleagues at Microsoft Research and Carnegie Mellon surveyed 319 knowledge workers about 936 real uses of AI at work. Higher confidence in the tool went with less critical thinking, and the thinking that remained shifted from producing to verifying. It is self-report, and people who think differently may use AI differently.

Kosmyna and colleagues at the MIT Media Lab ran the EEG study that launched the phrase cognitive debt: 54 participants writing essays with an LLM, a search engine, or unaided, with the LLM group showing the weakest brain connectivity. It is a preprint. Stankovic and three colleagues published a methodological critique arguing it is underpowered, that a repeated-measures design at those parameters would need around 159 participants, and that the search-engine group relied on an external tool and showed no impairment, with that comparison returning p = 1. The critique is itself unreviewed and presents no competing data. What it establishes is that the most-quoted study in this area rests on a contested pilot.

There is a pattern in that. The weaker a study's design, the closer its headline sits to the question people want answered, and the further it travels.

What the word smuggles in#

Stupid is a property of a person. Everything above is a property of a practice. That distinction changes who is responsible and what can be done.

Read as a claim about people, the evidence licenses nothing: there is no cohort getting dimmer. Read as a claim about practice, it licenses something quite sharp. A capability you stop exercising becomes a capability you cannot rely on, and the interval is shorter than anyone assumed. Liu's ten minutes is not skill decay, it is a within-session carry-over, and it still went in the direction nobody wanted. Budzyn's endoscopists had 27.6 years behind them. Experience did not insulate them.

This estate calls the accumulated version capability debt: the gap between what you can produce and what you could still do if the machine were switched off. It is invisible while the machine is on. The measurements above are all taken by switching it off, which is the only way anybody has found to see it.

The other thing the word hides is that the interface is doing most of the work. Bastani's three arms are the cleanest demonstration on record: the same technology, the same students, the same period, and the difference between a 17 per cent deficit and almost none was how the tool was built. That is a design question rather than a fact about human beings, which puts it in the hands of whoever chooses the tools.

Four things that follow for somebody using it daily#

Where this page is guessing#

The link between the withdrawal experiments and anything cumulative is unestablished. Every result above measures a person shortly after a period of assistance, over minutes, hours or a few months. Nothing follows the same people for years, so whether the effect compounds, plateaus or reverses is unknown, and Shen and Tamkin say in their own paper that whether an immediate quiz score predicts longer-term skill is a question they do not resolve. Budzyn is the longest-running observation here and it is observational. That paper also carries a correction, at Lancet Gastroenterol Hepatol 2025; 10: e12, whose content could not be read from here, so the 6.0 percentage points above is the figure as published and not a figure anybody here has checked against the amended paper.

Nor does anything measure the other side of the ledger. If attention freed from a task is spent on a harder one, the trade could be positive and no study here would detect it. The estate's position is that the withdrawal penalty is real, specific and the only thing anyone has measured, and that calling it stupidity overstates it in one direction while ignoring it understates it in the other.

Key sources

The long version of the second floor is at does using AI stop you learning and does AI weaken critical thinking. On what the withdrawal penalty accumulates into, cognitive debt and capability debt and deskilling. On the mechanism, cognitive offloading, the Google effect and productive struggle. On whether it comes back, can you regain a skill you have lost and how fast skills decay. For parents, should children use AI.

Evidence review · SS-2026-295 · Graded against the published rubric

Cite this page

Hirji, R. (2026). Is AI making us stupid?. The SuperSkills evidence base, SS-2026-295. https://thesuperskills.com/research/is-ai-making-us-stupid. Last reviewed 23 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

Is AI making us stupid?

Nobody has measured it. No study takes a person's general cognitive ability, gives them an AI assistant, and measures it again. What has been measured is unassisted performance after the tool is withdrawn, and that falls: high-school students with unrestricted GPT-4 scored 17 per cent lower than students who never had it once access was removed, and Polish endoscopists averaging 27.6 years of experience saw unassisted adenoma detection fall from 28.4 to 22.4 per cent after AI was introduced at their centres. Those are findings about a practice rather than about anybody's intelligence.

Does AI make you worse at thinking?

It makes you worse at what you have stopped doing, and the interval is shorter than anyone assumed. In randomised trials with 1,222 participants, Liu and colleagues found that assistance during a practice phase left people performing worse once it was withdrawn and more likely to give up, after roughly ten minutes of exposure. Their reading is that the first casualty is persistence rather than knowledge. Nothing shows a general decline in thinking ability.

Does AI make you lazy?

The closer description is that it removes the reason to persist. Liu and colleagues attribute the withdrawal penalty they measured to AI conditioning people to expect immediate answers, which denies them the experience of working through a problem alone. The tell is reaching for the tool earlier in a problem than you used to, rather than failing to remember something. The design of the tool matters more than the character of the user: a hints-only tutor produced better learning than unrestricted access in the same experiment, on the same students.

Is AI bad for your brain?

The study usually cited for this is the MIT Media Lab EEG work of Kosmyna and colleagues, which found the weakest brain connectivity in the group writing essays with a language model. It has 54 participants, it is a preprint, and a published critique by Stankovic and three colleagues argues it is underpowered and points out that the search-engine group also relied on an external tool and showed no impairment, with that comparison returning p = 1. The critique is itself unreviewed. Treat claims of proof in either direction with suspicion.

Are IQ scores falling because of AI?

No. The clearest population evidence of falling cognitive test scores is Bratsberg and Rogeberg's analysis of Norwegian military conscription records for birth cohorts from 1962 to 1991, covering 736,808 scored men. Mean IQ peaked with the 1975 cohort at 102.3 and fell to 99.4 by 1989. The 1991 cohort sat that test around 2009, so the entire decline predates generative AI. The authors recovered the fall from within-family variation, which rules out explanations based on who was having children, and they state that they cannot identify which environmental factor is responsible.

In this hub

Thinking, learning and capability

What sustained AI use does to thinking, and how capability is built and kept.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory for CEOs and boards  ·  Enquire