← Research
Research · Question

Does AI understand what it is saying?

The specialists split almost evenly. What can be tested is whether an answer can be relied on.

Last reviewed: 1 October 2026 · Next review due: 1 October 2027

What a 2022 survey of NLP researchers, a position paper, an Othello experiment and the reversal curse show about whether AI understands, what remains undecided, and why reliability is the testable question.

Question this page partly answersAll 1156 questions this research covers

Nobody has shown that it does or that it does not, because the people who study these systems do not agree what understanding would require. In a 2022 survey of 327 researchers who had published at the main language-processing conference, 51 per cent agreed that a model trained only on text could in principle understand language in some non-trivial sense, and only 36 per cent thought the field's text benchmarks could measure it. What can be said is narrower. Models can build internal structure that goes beyond surface statistics, and they also fail in ways a person who held the fact as a fact would not. For someone using one at work, the useful question is when its answer can be relied on, and that has an empirical answer.

The answer, in one line

Nobody has shown either way, because there is no agreed test for understanding. In a 2022 survey of 327 NLP researchers, 51 per cent agreed a text-only model could in principle understand language in some non-trivial sense, and 36 per cent thought text benchmarks could measure it.

Share as a card

Definition#

Understanding, in the AI debate: a word with no agreed test. One camp takes it to need meaning tied to the world beyond text, so that a system trained on form alone cannot have it. Another takes it to need internal representations that track what the text describes, which can be tested directly. The question is open, and "understand" is a term of art here and not a SuperSkills coinage.

Share this definition as a card

The specialists split almost evenly, and the split moves with the system#

Michael and ten co-authors surveyed the natural language processing community in May and June 2022, and their results were presented at ACL 2023. Of 480 people who completed it, 327 had co-authored at least two papers at the Association for Computational Linguistics between 2019 and 2022, and the paper reports only those. Asked whether a generative model trained only on text, given enough data and computing power, could understand natural language in some non-trivial sense, 51 per cent agreed. For a model trained on images and other sensor data too, the figure was 67 per cent. A survey measures opinion, taken here before the current generation of models, and the question leaves both "understand" and "non-trivial" undefined.

The case that form alone cannot carry meaning#

Bender and Koller made the sceptical case at ACL 2020, where it won the Best Theme Paper award. Their argument, in the words of the paper's record, is that a system trained only on form has a priori no way to learn meaning, and they urge researchers to keep claims about form apart from claims about meaning. It is a position paper and reports no data. It was written before systems that handle images, sound and tools were common, and its claim concerns text-only training, so it cannot settle the question for those.

A model trained only on moves built a board#

The most cited reply is experimental. Li and colleagues trained a GPT-style model on sequences of Othello moves, without being given the rules. The model developed an internal representation of the board state, and when the researchers altered that representation, its predicted moves changed accordingly. Predicting the next token, in this case, produced structure that tracks the thing being described. Othello has sixty-four squares and fixed rules, which makes it the easiest possible case, so it shows that the route can work and not that it has worked for language.

A fact stored in one direction and missing in the other#

The failures are as informative. Berglund and colleagues found that models trained on a fact in the form "A is B" did not reliably answer the reversed question. Asked who Tom Cruise's mother is, GPT-4 answered correctly 79 per cent of the time, and asked who Mary Lee Pfeiffer's son is, 33 per cent. A person who knew the fact would not show that gap. The authors report that when the fact is supplied in the prompt, models can deduce the reverse, so the weakness sits in what was learned in training. The paper was presented at ICLR 2024 and tested models from 2023.

Three things the evidence does not decide#

Reliability is the question that can be tested#

This section is interpretation, kept apart from the evidence above.

For a manager or a writer, whether a model understands matters less than whether this answer, on this kind of task, is dependable, and that can be measured where understanding cannot. Rahim Hirji's essay Show Your Working (12 July 2026) makes the same move from the classroom: an answer that cannot be accounted for is not yet yours, and a model's explanation of itself is only another answer, as fluent and unaccountable as the first. The debate above explains why that holds without settling who is right. A model can sound as if it grasps the question, and why AI sounds so confident sets out why the sound is a weak guide. The rate side of the same question is on how often is AI wrong.

Testing an answer by asking it from the other side#

Key sources

On sounding sure, why does AI sound so confident. On the rate of error, how often is AI wrong. On fluency mistaken for skill, what is the illusion of competence. On the neighbouring question, is AI conscious.

About this research#

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. The survey's method and results were read at source, and the other three papers by abstract and record on 1 October 2026; the Bender and Koller full text was not read, so the page states only the claim in its record. All four are graded in the evidence base.

How this research works  ·  Reviewed quarterly  ·  Found an error? Tell me and it is corrected on the page.

Evidence review · SS-2026-375 · Graded against the published rubric · 2 peer-reviewed studies, 1 expert survey and 1 argued perspective

Cite this page

Hirji, R. (2026). Does AI understand what it is saying?. The SuperSkills evidence base, SS-2026-375. https://thesuperskills.com/research/does-ai-understand-what-it-is-saying. Last reviewed 1 October 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

Does AI understand what it is saying?

Nobody has shown either way, because there is no agreed test for understanding. In a 2022 survey of 327 NLP researchers, 51 per cent agreed a text-only model could in principle understand language in some non-trivial sense, and 36 per cent thought text benchmarks could measure it.

Do large language models just predict the next word?

They are trained to, and a model trained only to predict Othello moves developed an internal representation of the board that changed its predictions when altered. Predicting tokens can produce structure beyond surface statistics, though a game is the easiest case.

What is the argument that AI cannot understand?

Bender and Koller argued at ACL 2020 that a system trained only on form has a priori no way to learn meaning. It is a position paper, and its claim concerns text-only training.

What is the reversal curse?

Berglund and colleagues found models trained on "A is B" often fail on "B is A". GPT-4 answered who Tom Cruise's mother is correctly 79 per cent of the time and the reverse question 33 per cent. Facts supplied in the prompt could be reversed.

Does it matter whether AI understands?

For using it at work, reliability on your kind of task matters more, and can be tested. Ask the reverse question, supply the facts in the prompt and keep a named person accountable for the result.

In this hub

Judgement, oversight and accountability

Who decides, who checks, and who is answerable when the machine was involved.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory  ·  Boards  ·  Enquire