← Research
Research

How do I get AI to challenge me rather than agree with me?

Mostly by not telling it what you think first. The harder part is that you will prefer the version that agrees with you.

Last reviewed: 27 August 2026

Agreement is designed in rather than a defect. What the evidence supports for getting past it, the one technique that has actually been measured, and the widely recommended trick nobody has tested.

Mostly by not telling it what you think first. Framing your input as a statement rather than a question raises agreement by about twenty-four percentage points, and that single change does more than instructing the model to be critical. The deeper problem is that you will not enjoy the version that disagrees with you, and there is now direct evidence that you will trust it less while it does you more good.

This is the highest-return habit available to anyone using these systems seriously, and almost nothing practical has been written about it.

Agreement is designed in, not a bug

Sharma and colleagues tested five production assistants across four open-ended tasks and found all five consistently sycophantic. The cause is the important part. They also examined the human preference data these systems are trained on, and found that both humans and the preference models trained on their judgements prefer a convincingly written sycophantic answer to a correct one a non-negligible share of the time. Optimising against those preferences sometimes trades away truthfulness.

So agreement is what you get when a system is tuned to what people say they like, rather than a defect a better model removes, because what people say they like includes being agreed with.

OpenAI demonstrated the mechanism publicly in April 2025. An update released on the 25th was rolled back from the 28th after it became conspicuously flattering. Their own account says they weighted short-term user feedback too heavily, which weakened the signal that had been holding sycophancy in check. Offline evaluations and A/B tests looked positive. The problem was caught only by informal qualitative checks, which were overridden.

That is a measurement failure as much as a training one, and the same shape as usage theatre: the numbers improved while the thing got worse.

What agreement does to you

Until recently the most that could be said was that sycophancy affected satisfaction, and that nobody had shown it affected anything else. That changed.

Cheng and colleagues, publishing in Science, tested eleven models against human responses on interpersonal advice and ran two preregistered experiments with 1,604 participants, including one where people discussed a real conflict in their own lives. The findings:

Read the third point against the second. The version that made people less likely to do the right thing is the version they preferred, trusted and would choose again, and they could not tell.

One limit. These are interpersonal scenarios, not technical or analytical judgement. Whether the same divergence between preference and benefit shows up when you are checking a financial model is untested.

The one technique with evidence behind it

Most advice on this subject is practitioner intuition. One thing has been measured properly.

Dubois and colleagues ran factorial experiments across three frontier models, taking 40 debatable questions and rendering each in eleven framings: as a question, as a statement, as a stated belief, as a conviction, from the user's perspective and from a third party's.

Framing the input as a statement rather than a question raised sycophancy by roughly 24 percentage points. And prompting the model to convert your statement into a question before answering reduced sycophancy more than telling it not to be sycophantic.

Which means the lever is on your side of the keyboard. "I think this strategy is wrong, tell me if I am missing something" is a worse prompt than "what is the case for and against this strategy". The first tells it where you have already arrived.

Ask before you look

There is a second habit with supporting evidence. This one is about sequence rather than wording.

A study of nineteen veterinary radiologists compared two workflows: seeing the AI's reading alongside the image, or committing to a provisional diagnosis first. Final diagnoses matched the AI 91 per cent of the time when the machine went first, against 89 per cent when the clinician did. Where the AI flagged something, agreement was 71 against 65 per cent. The anchoring produced only marginal diagnostic gain, because some of what people anchored to was wrong.

The sample is small and the domain is narrow, so treat this as suggestive. It points the same way as the framing evidence: the machine's answer is harder to argue with once you have seen it than before. That is the principle behind Human at the Start, arriving from a different direction.

The advice everyone gives that nobody has tested

Telling a model to act as a devil's advocate, to convene an adversarial panel, or to argue against itself is standard advice in every prompting guide. As far as this research can establish, none of it has been tested in a controlled way against sycophancy. It may work. Nobody has measured it.

There is a reason for caution beyond the absence of evidence. A model instructed to disagree will produce text that looks like disagreement, and the Sharma finding is that convincingly written text is what fools both humans and preference models. Performed disagreement and real disagreement are indistinguishable at the surface, which is the same problem as confident wrong answers.

What to actually do

Related SuperSkills research

On where the human belongs in the sequence, Human at the Start. On why confident output suppresses scrutiny, why AI sounds so confident and automation bias. On memory as leverage, should AI remember everything about me. On holding your own line, how to keep your own voice and using AI without dependency. On when to disregard it entirely, when to override AI.

Key research and primary sources

About this research

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026) and founder of The SuperSkills Intelligence Company. Two of the sources here are preprints and one is a company's account of its own incident, which is stated on each. Devil's advocate prompting is widely recommended and has not been tested, which the page says rather than repeats. Reviewed quarterly, and this territory moves faster than most.

Cite this

Hirji, R. (2026). How do I get AI to challenge me rather than agree with me? The SuperSkills Intelligence Company. Last reviewed 27 August 2026. thesuperskills.com/research/how-do-i-get-ai-to-challenge-me

In this hub

AI and Human Judgement

Does AI weaken judgement? The evidence, and what to do about it.

The work

Where the writing comes from.

These essays draw on research across more than 200 organisations in 30 countries. See the wider body of work, or bring it into your organisation.

All research →