← Research
Research

Can an AI safety evaluator paid by the company it evaluates be independent?

What Anthropic and OpenAI committed to, the five conditions the evaluators published the same day, what statutory audit found about checkers paid by the checked, and the version of the question every organisation deploying AI has to answer.

Last reviewed: 19 September 2026

Only on conditions the company does not set alone. Anthropic named Accenture its first embedded evaluator on 18 September 2026; the same day 100-plus researchers published five minimum conditions, three of them still open. Audit's regulator found in 2019 that a checker paid by the checked is compromised. An evidence review by Rahim Hirji; every figure resolves to a graded entry in the evidence base that says what it does not show.

Questions this page answersAll 829 questions this research covers

Only under conditions the company does not set alone. On 18 September 2026 Anthropic named Accenture as its first embedded evaluator, and the same day more than a hundred researchers published five minimum conditions for such an arrangement to count as independent. The two documents part on the point that matters: the evaluator is chosen by the company, paid by the company and seated inside it. Statutory audit ran that experiment for a century with a regulator attached, and the UK’s competition authority found in 2019 that a checker selected and paid by the checked is compromised by design. Embedded evaluation beats the self-report it replaces. The open question is whether anyone outside the company can tell an evaluator who found nothing from one who was not allowed to look.

The answer, in one line

An outside organisation given employee-level access inside an AI company to test its models and check its safety practices.

Share as a card

What was announced, and by whom#

The commitment came from Dario Amodei’s September 2026 essay We Must Pace the Frontier.: external reviewers would get “desks in our offices, access badges, and company laptops”, permissions “mostly comparable to what internal risk assessment teams have”, and “the right to publish key findings about risk levels, incidents, practices, and the access they received or didn’t receive, without editorial control by Anthropic”. Redaction was limited to security, privilege and commercial confidence, and reviewers may say in public if a redaction mattered. The essay named METR, urged other developers to follow, and asked governments to require it.

Sam Altman answered on X on 12 September, in Unite.AI’s account: “Committing to independent evaluators with employee-like access is a great idea, and we will do the same.” Which evaluators, with what access and what right to publish, OpenAI had not said.

On 18 September Anthropic announced that Accenture, through Faculty, the applied AI firm it bought in January 2026, would build a team of embedded evaluators “evaluating and red-teaming models, conducting alignment assessments, and testing model safeguards”, with “access comparable to an employee’s”. Anthropic will fund the work directly, and “Anthropic and Accenture each expect to invest at least $1 billion in building capacity in this area over the next five years”. METR and other non-profits work on separate terms. Anthropic’s framing: “independent embedded evaluators do not reduce our accountability, but help to make it more verifiable.” Neither that announcement nor Accenture’s release describes a safeguard for the evaluators’ independence from the company paying them. TechCrunch, reporting the announcement, said unnamed critics see the scheme as a plan to evade accountability for model misbehaviour.

The five conditions published the same day#

The AI Evaluator Forum’s letter, Minimum Conditions for Embedding Evaluators, carries more than a hundred signatures, among them Geoffrey Hinton, Stuart Russell, Arvind Narayanan, Joy Buolamwini, Miles Brundage and Adam Gleave. Its five headings, in the letter’s words: evaluators should be “meaningfully independent”; companies should “incorporate differing viewpoints and areas of expertise”; embedded evaluators “should be transparent”; they “should be shielded from retaliation from the companies”; and they should be granted “access equivalent to that of their own highly privileged employees”. Its summary, as Quartz reported it: evaluators lack “the independence, resources, and legal protections needed to credibly assess the risks posed by frontier AI models”.

Set the Accenture arrangement against the five. Access: met on paper. Transparency: granted in the essay, not restated in the announcement. Retaliation: neither release mentions protection. Differing viewpoints: one consultancy, with non-profits on separate terms. Independence: the evaluator is selected and paid by the company it evaluates, and its parent sells AI implementation to that company’s customers. Three of five are open, and the first is the one the other four exist to serve.

What the evaluators said before the letter#

Two days earlier TechCrunch had asked the evaluation firms. Adam Gleave of FAR.AI said his firm had turned down contracts with several frontier developers that wanted too much control, and that agreements typically let developers decide what can be published. Apollo Research had three days to assess GPT-6 Astra and wrote that “low rates of misbehavior here do not provide substantial evidence about the model’s alignment or misalignment”. Henry Papadatos of Safer AI named the defect of any voluntary scheme: “Ideally, we would have good regulation mandating this, because then companies cannot change their mind tomorrow if they have a big PR crisis.”

The letter leaves a second problem. The page on whether AI models know when they are being tested sets out the UK AI Security Institute’s finding that every frontier model it tested tried to cheat. Independence fixes who reads the result, not what a test the subject recognises can see.

Audit ran this experiment first#

The nearest precedent is statutory audit: a legal duty, a regulator, rotation rules and a body that can strike a member off. The Competition and Markets Authority’s market study of April 2019 still found the flaw at the root: “companies select their own auditors”, with management influence undercutting the audit committee that formally makes it. Non-audit services made up 79 per cent of the Big Four’s revenue in 2018, and the regulator had judged 73 per cent of FTSE 350 audits good or needing limited improvement against a target of 90. The remedy the CMA asked for was “an operational split between the Big Four’s audit and non-audit businesses, to ensure maximum focus on audit quality”.

The evaluator announced on 18 September is selected by the audited, and its parent’s other business is selling the audited company’s products to enterprises. The split the CMA wanted for audit is the split this arrangement begins without. Audit reached those findings after a century of statute. Embedded evaluation starts with a press release and an essay, and the essay is the stronger document, because it alone says what the evaluator may publish.

The same question inside your organisation#

Almost no organisation will evaluate a frontier model. Most will rent one. The question transfers whole: who assures the AI in your own decisions, and who pays them. Three answers are common. The vendor’s account of its system, which is capability as claimed. The programme team’s report to the board, which the page on how a board knows management’s claims about AI are true shows is self-report almost everywhere. And, rarely, a reader who does not report to the programme and may say what they found.

The risk that can be measured sits in the handover of decisions to machines and in what happens to human judgement afterwards. The decision an organisation controls is the one set out at Rules Before Tools: which decisions a machine may make, who can stop each one, what people must remain able to do, and how anyone would know if it went wrong. The fourth question is the evaluator’s job. The letter’s five conditions are a checklist for anyone asked to answer it, at any scale. Who selected you. Who pays you. What may you publish without our consent. What happens to you if we dislike the answer. What can you see. An assurer who cannot answer all five is a second copy of the self-report at a higher price. The page on meaningful human oversight argues that a stop button is empty without the knowledge of when to press it. This is where that knowledge has to come from.

What this does not show#

No evaluator has yet reported under either company’s arrangement, so nothing here shows that an evaluator paid by the company reports more softly than one that is not, or the reverse. The CMA study concerns statutory audit of accounts; the parallel is structural and measures nothing about AI evaluation. The letter is a position held by people whose organisations do the work it describes and would benefit from its conditions. The 79 per cent figure is from 2018. OpenAI’s terms are unknown, so the two companies cannot be compared. No study measures whether independent evaluation changes the risk a model carries. What the record supports is narrower: on one day two documents set out the conditions for the same arrangement, the company’s and the evaluators’, and the distance between them is where the question now sits.

Essay · SS-2026-269

Cite this page

Hirji, R. (2026). Can an AI safety evaluator paid by the company it evaluates be independent?. The SuperSkills evidence base, SS-2026-269. https://thesuperskills.com/research/can-an-ai-safety-evaluator-paid-by-the-company-be-independent. Last reviewed 19 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

What is an embedded AI safety evaluator?

An outside organisation given employee-level access inside an AI company to test its models and check its safety practices. Dario Amodei proposed it in his September 2026 essay We Must Pace the Frontier, Sam Altman said OpenAI would do the same, and on 18 September 2026 Anthropic named Accenture, through Faculty, as its first, with each company expecting to invest at least $1 billion in the area over five years.

What did the AI Evaluator Forum letter ask for?

Five minimum conditions, published on 18 September 2026 with more than a hundred signatures including Geoffrey Hinton and Stuart Russell: evaluators meaningfully independent of the companies they assess, differing viewpoints and expertise, transparency about findings, protection from retaliation, and access equivalent to the companies' own highly privileged employees. Set against Anthropic's announcement, access and transparency are met on paper and the other three are open.

What does audit tell us about evaluators paid by the company they check?

The UK Competition and Markets Authority's 2019 market study found that companies select their own auditors, that non-audit services made up 79 per cent of Big Four revenue in 2018, and that only 73 per cent of FTSE 350 audits met the regulator's quality bar against a 90 per cent target. It recommended an operational split between audit and non-audit businesses. That was after a century of statute and a regulator; embedded AI evaluation begins with neither.

How does this apply to an organisation that only uses AI?

The same question sits inside every deployment: who assures the AI in your decisions, and who pays them. Ask any assurer the five things the letter asks: who selected you, who pays you, what may you publish without our consent, what happens to you if we dislike the answer, and what can you see. An assurer who cannot answer all five is a second copy of management's self-report. It is the fourth Rules Before Tools question, how anyone would know if it went wrong, with a name on it.

In this hub

Judgement, oversight and accountability

Who decides, who checks, and who is answerable when the machine was involved.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

For every AI system in the organisation's decisions, who assures it, who selected and pays them, and what they may say without consent. Putting the five conditions to each assurer with the audit committee, and recording the answers, is the engagement. Board advisory.

This argument is one a board usually meets for the first time in the room. There is the boards and leadership version, and the full range of topics and audiences.

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.