Nobody has shown that it does or that it does not, because the people who study these systems do not agree what understanding would require. In a 2022 survey of 327 researchers who had published at the main language-processing conference, 51 per cent agreed that a model trained only on text could in principle understand language in some non-trivial sense, and only 36 per cent thought the field's text benchmarks could measure it. What can be said is narrower. Models can build internal structure that goes beyond surface statistics, and they also fail in ways a person who held the fact as a fact would not. For someone using one at work, the useful question is when its answer can be relied on, and that has an empirical answer.
The answer, in one line
Nobody has shown either way, because there is no agreed test for understanding. In a 2022 survey of 327 NLP researchers, 51 per cent agreed a text-only model could in principle understand language in some non-trivial sense, and 36 per cent thought text benchmarks could measure it.
Definition#
Understanding, in the AI debate: a word with no agreed test. One camp takes it to need meaning tied to the world beyond text, so that a system trained on form alone cannot have it. Another takes it to need internal representations that track what the text describes, which can be tested directly. The question is open, and "understand" is a term of art here and not a SuperSkills coinage.
The specialists split almost evenly, and the split moves with the system#
Michael and ten co-authors surveyed the natural language processing community in May and June 2022, and their results were presented at ACL 2023. Of 480 people who completed it, 327 had co-authored at least two papers at the Association for Computational Linguistics between 2019 and 2022, and the paper reports only those. Asked whether a generative model trained only on text, given enough data and computing power, could understand natural language in some non-trivial sense, 51 per cent agreed. For a model trained on images and other sensor data too, the figure was 67 per cent. A survey measures opinion, taken here before the current generation of models, and the question leaves both "understand" and "non-trivial" undefined.
The case that form alone cannot carry meaning#
Bender and Koller made the sceptical case at ACL 2020, where it won the Best Theme Paper award. Their argument, in the words of the paper's record, is that a system trained only on form has a priori no way to learn meaning, and they urge researchers to keep claims about form apart from claims about meaning. It is a position paper and reports no data. It was written before systems that handle images, sound and tools were common, and its claim concerns text-only training, so it cannot settle the question for those.
A model trained only on moves built a board#
The most cited reply is experimental. Li and colleagues trained a GPT-style model on sequences of Othello moves, without being given the rules. The model developed an internal representation of the board state, and when the researchers altered that representation, its predicted moves changed accordingly. Predicting the next token, in this case, produced structure that tracks the thing being described. Othello has sixty-four squares and fixed rules, which makes it the easiest possible case, so it shows that the route can work and not that it has worked for language.
A fact stored in one direction and missing in the other#
The failures are as informative. Berglund and colleagues found that models trained on a fact in the form "A is B" did not reliably answer the reversed question. Asked who Tom Cruise's mother is, GPT-4 answered correctly 79 per cent of the time, and asked who Mary Lee Pfeiffer's son is, 33 per cent. A person who knew the fact would not show that gap. The authors report that when the fact is supplied in the prompt, models can deduce the reverse, so the weakness sits in what was learned in training. The paper was presented at ICLR 2024 and tested models from 2023.
Three things the evidence does not decide#
- Whether behaviour is enough. If understanding means reliable use of meaning across contexts, failures like the reversal curse count against it. If it means grounded experience, no behavioural test can settle it.
- Whether today's models differ. Every study above predates October 2026 by at least a year, and none tested the reasoning and tool-using models in use today.
- Whether the word helps. Consciousness is a separate question from understanding, and is AI conscious covers the evidence there.
Reliability is the question that can be tested#
This section is interpretation, kept apart from the evidence above.
For a manager or a writer, whether a model understands matters less than whether this answer, on this kind of task, is dependable, and that can be measured where understanding cannot. Rahim Hirji's essay Show Your Working (12 July 2026) makes the same move from the classroom: an answer that cannot be accounted for is not yet yours, and a model's explanation of itself is only another answer, as fluent and unaccountable as the first. The debate above explains why that holds without settling who is right. A model can sound as if it grasps the question, and why AI sounds so confident sets out why the sound is a weak guide. The rate side of the same question is on how often is AI wrong.
Testing an answer by asking it from the other side#
- Ask the reverse question. If a model tells you a fact about a person, company or document, ask the converse. The reversal result suggests the two directions can come apart. The check comes from that finding and nobody has validated it as a test.
- Supply the facts in the prompt. Models handled reversal when the fact was in front of them, so working from a document you have pasted in is safer than relying on recall.
- Keep the judgement about fit with a person. A model can be fluent about a decision without carrying its consequences. Decide who owns the answer before it is used.
Key sources
- Michael, J. et al. (2022). What Do NLP Researchers Believe? Results of the NLP Community Metasurvey. arXiv:2208.12852; ACL 2023. Graded entry.
- Bender, E. M. and Koller, A. (2020). Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data. ACL 2020, 5185 to 5198. Graded entry.
- Li, K. et al. (2023). Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task. ICLR 2023. Graded entry.
- Berglund, L. et al. (2024). The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A". ICLR 2024. Graded entry.
Related SuperSkills research#
On sounding sure, why does AI sound so confident. On the rate of error, how often is AI wrong. On fluency mistaken for skill, what is the illusion of competence. On the neighbouring question, is AI conscious.
About this research#
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. The survey's method and results were read at source, and the other three papers by abstract and record on 1 October 2026; the Bender and Koller full text was not read, so the page states only the claim in its record. All four are graded in the evidence base.
Evidence review · SS-2026-375 · Graded against the published rubric · 2 peer-reviewed studies, 1 expert survey and 1 argued perspective
Hirji, R. (2026). Does AI understand what it is saying?. The SuperSkills evidence base, SS-2026-375. https://thesuperskills.com/research/does-ai-understand-what-it-is-saying. Last reviewed 1 October 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work