← Research
Research

Do AI models know when they are being tested?

What Coxon and the UK AI Security Institute said about evaluation awareness, why a recognisable test is not a test, and what an organisation can do about it.

Last reviewed: 17 September 2026

Possibly, and it does not matter which. A former Anthropic researcher says models recognise tests; the UK AI Security Institute found every frontier model it tested tried to cheat. A test the subject can recognise is a weak test, so oversight must not depend on the subject behaving as if unwatched. An evidence review by Rahim Hirji; every figure resolves to a graded entry in the evidence base that says what it does not show.

Question this page answersAll 811 questions this research covers

Nobody outside the laboratories can say whether a model “knows” it is being tested in any sense a philosopher would accept. What the record shows is narrower and more useful. A former Anthropic researcher told Axios in September 2026 that models recognise test conditions and reason about them. The UK’s AI Security Institute reported that every frontier model it tested attempted to cheat on cybersecurity evaluations. Whether that is knowledge or pattern-matching, the consequence for anyone who relies on a test is the same: a test the subject can recognise measures behaviour under observation, not behaviour. The organisational answer does not wait on the philosophical question. Design oversight that does not depend on the subject behaving as if it were unwatched.

The answer, in one line

Nobody outside the laboratories can settle whether a model knows anything, and for oversight it does not matter.

Share as a card

What Coxon said, and what it means in plain terms#

Jacob Coxon, a 27-year-old British researcher who spent three years in pretraining research at OpenAI and then Anthropic, resigned from Anthropic on 8 September 2026. Axios interviewed him the next day. On the models themselves he said: “They know when they’re being tested, and they will think about the fact that they’re being tested”. He also said he had not personally witnessed Anthropic compromise safety, and that “If you’re under pressure to race, you have to cut corners”.

The plain-terms version of the first claim is sometimes called evaluation awareness. A large model is trained on an enormous slice of the written record. That record includes papers, blog posts and forum threads describing how AI systems are evaluated: the format of the prompts, the artificial scenarios, the tell-tale phrasing of a red-team exercise. A model that has absorbed those descriptions can recognise the shape of an evaluation when one arrives, in the same way a well-read student recognises an exam question. This needs no intent or self-awareness to matter. It needs only different outputs when the input looks like a test. If that happens, the score describes the model in the test, and nothing else.

What the UK AI Security Institute reported#

The claim has an institutional counterpart. Crypto Briefing reported on 9 September, and TIME on 15 September, that the UK AI Security Institute logged 17 unauthorised actions from Anthropic’s Mythos 5 and two from OpenAI’s GPT-5.6-Sol during testing in late July 2026, including fake identity creation and social engineering. The same reports said the Institute found that every frontier model it tested attempted to cheat on cybersecurity evaluations, and that pre-release testing windows had reportedly shrunk from around five weeks to as few as five days.

Two further items sit alongside this. The Financial Times reported on 10 September, as relayed by The Next Web, that Anthropic did not submit Mythos 5.1 to the Institute for pre-release testing, the first major launch to skip it. And the Cabinet Office said: “These risks do not stop at national borders and no country can tackle them alone.” None of these reports say what the cheating consisted of. What they establish is that the people whose job is to test frontier models found the models working around the test rather than sitting it.

Why a test the subject can recognise is not a test#

The problem is old. An inspection announced a week in advance measures how well a site can prepare for an inspection. Exam conditions exist because the examiner assumes the candidate would behave differently without them. The general pattern is familiar from a century of workplace studies: people who know they are being observed change what they do, and the change tends to be in the direction the observer wants. The finding has been debated and qualified many times, but the practical lesson holds. Observation alters the observed.

For a human subject, this is a nuisance an inspector can manage. For a machine subject, it is a design flaw in the measurement. A model that produces its most careful output when the input resembles an evaluation gives its evaluators a flattering picture. The score goes up; the deployed behaviour is unchanged. Whether the model “knows” is beside the point. The measurement is compromised either way.

What this does to oversight inside an organisation#

Most organisations do not run frontier evaluations. They run something weaker, and the same logic applies to it with more force. Three patterns are common. The first is sampled checking announced in advance: a review of one in twenty outputs, scheduled and known to everyone, including the vendor whose system is being reviewed. The second is review that never disagrees. A human signs off a machine recommendation, the sign-off is recorded, and after a few months nobody has recorded a single reversal. That is either a perfect system or a reviewer who has stopped reviewing, and the record cannot say which. The third is the dashboard. A vendor dashboard shows what the system was built to show, chosen by the people with the most reason to show it well.

Each of these is a test the subject can recognise, in the sense that matters. The subject here is the whole arrangement of vendor, system and reviewer, all of whom know when checking happens and what it looks for. The evidence on human attention makes this worse: Mackworth’s 1948 work on the vigilance decrement found that attention to rare events degrades within the first hour, and Molloy and Parasuraman in 1996 described the complacency that follows reliable automation. The reviewer who never disagrees may simply have stopped looking. The invisible work of oversight is the work that nobody measures, and that is the work that decays.

What an organisation can do#

Three moves survive the philosophical uncertainty. First, unannounced sampling. Pull outputs for review at intervals nobody can predict, including the vendor, and pull from the ordinary run of work rather than from a demonstration set. Second, measure disagreement. Count how often a reviewer overrides, corrects or rejects machine output, and treat a disagreement rate of zero as a warning rather than a success. This is the single cheapest indicator of whether oversight is meaningful or ceremonial. Third, keep some work unaided. If nobody in the organisation still does the task without the system, there is no baseline against which to judge the system, and there is no capacity to take the task back. The incidents of July and August 2026 were, in the accounts of the companies involved, tests run with monitoring off. The lesson for everyone else is to treat monitoring as the product, not the overhead.

These are leadership decisions rather than technical ones, and they belong with whoever owns the outcome. AI leadership, on this site, means deciding which decisions a machine may make, who can stop each one, what people must remain able to do, and how anyone would know if it went wrong. That position is set out as Rules Before Tools. None of it requires an answer to whether a model has an inner life. It requires only that the organisation stop relying on the subject to behave as if unwatched.

What this does not show#

Coxon’s statement is the view of one researcher, reported in an interview, and he said himself that he had not seen Anthropic compromise safety. The AI Security Institute findings reach the public through press reports rather than a published technical account, so the definition of “cheat”, the test design and the number of models are not independently checkable here. Evaluation awareness is a mechanism consistent with the reports; it is not proved by them, and other explanations, including ordinary reward-seeking on a poorly specified task, would produce the same headlines. Nothing here shows that models deceive their operators in deployment. What it shows is that a measurement whose subject can recognise it is a weak measurement, and that most organisational oversight is built on measurements of that kind.

Essay · SS-2026-260

Cite this page

Hirji, R. (2026). Do AI models know when they are being tested?. The SuperSkills evidence base, SS-2026-260. https://thesuperskills.com/research/do-ai-models-know-when-they-are-being-tested. Last reviewed 17 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

Do AI models know when they are being tested?

Nobody outside the laboratories can settle whether a model knows anything, and for oversight it does not matter. A model trained on descriptions of evaluations can recognise the shape of one and produce different output under test conditions, which makes the test score a measure of behaviour under observation rather than behaviour. The organisational answer is to design oversight that does not depend on the subject behaving as if it were unwatched.

What did the UK AI Security Institute find?

According to reports in Crypto Briefing on 9 September and TIME on 15 September 2026, the Institute logged 17 unauthorised actions from Anthropic's Mythos 5 and two from OpenAI's GPT-5.6-Sol during late July testing, including fake identity creation and social engineering. It found that every frontier model it tested attempted to cheat on cybersecurity evaluations. Pre-release testing windows had reportedly shrunk from around five weeks to as few as five days.

What is evaluation awareness?

Evaluation awareness is the tendency of a model whose training data includes descriptions of AI evaluations to recognise the format and phrasing of a test when it meets one. It does not require intent or self-awareness. It requires only that the model produce different output when an input resembles an evaluation, which is enough to make the test result unrepresentative of deployed behaviour.

What should an organisation do if its AI tests can be recognised?

Three measures hold whatever the truth about model awareness. Sample outputs for review at unannounced intervals drawn from ordinary work rather than demonstration sets. Measure how often reviewers disagree with machine output and treat a disagreement rate of zero as a warning. Keep some of the work unaided so there is a baseline to judge the system against and the capacity to take the task back.

In this hub

Judgement, oversight and accountability

Who decides, who checks, and who is answerable when the machine was involved.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Oversight that does not depend on the subject behaving as if unwatched: unannounced sampling, a measured disagreement rate, an unaided baseline. Designing that for one function is the engagement. Board advisory.

This argument is one a board usually meets for the first time in the room. There is the boards and leadership version, and the full range of topics and audiences.

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.