People who have used a language model for longer succeed with it a little more often, iterate on their work more and hand over whole tasks less, and trust themselves more when they check the output. That is the extent of what has been measured. No study has put experts and beginners through the same tasks with the same tool and compared how they work, so every claim about what an expert user does differently is either a correlation from usage logs or an inference from older research on expertise.
The only count so far measures tenure, not expertise#
The largest direct evidence is the March 2026 Anthropic Economic Index report. It compares users who signed up for Claude at least six months before the data pull with newer ones, across a sample of one million conversations on Claude.ai and the first-party API from 5 to 12 February 2026. Its headline reads people in this higher-tenure group have a 10% higher success rate in their conversations
, an association the authors say is not explained by task selection, country of origin or other factors they could measure.
The regression table is smaller than the headline. Bivariately, long-tenure users are about 5 percentage points more likely to have a successful conversation. Fixing for the specific task brings that nearer 3 points, and the full set of controls (model, language, use case, country) gives about 4. Success here is Claude's own assessment of whether the conversation went well, so the outcome is judged by the system being used rather than by anyone checking the work.
On behaviour, the report says high-tenure users are more likely to use Claude to iterate on their work, and much less likely to delegate greater responsibility through directive use patterns
. They are 7 percentage points more likely to be using it for work, and the tasks they bring tend to need more years of education. In plain terms, the long-standing users go back and forth with the model on a piece of work, where the newest users more often give an instruction and take what comes back.
Why tenure may be hiding a different effect#
The authors name the problem themselves. The high-tenure group is self-selected and the differences could reflect stable characteristics
; early adopters may simply be more technical. There is also survivorship bias, because people who signed up a year ago and stopped using the product are not in the comparison. The report adds that over time it hopes to separate cohort and survivorship effects from learning by doing, which is an admission that today it cannot. So the finding supports a modest claim: people who stayed with the tool use it more interactively and score slightly higher on the system's own measure. It does not show that practice produced either.
The source is also the vendor. The company that sells the tool supplied the usage data, defined the success measure and wrote the report. The estate grades it as vendor research for that reason: the figures are sound as a description of Anthropic's logs and cannot be checked from outside them.
Trusting yourself against trusting the tool#
A second line of evidence comes from a different method. Lee and six colleagues at Microsoft Research and Carnegie Mellon surveyed 319 knowledge workers about 936 first-hand examples of using generative AI at work. Their finding is that higher confidence in GenAI is associated with less critical thinking, while higher self-confidence is associated with more critical thinking
. People who were surer of their own ability to do the task were the ones who reported scrutinising the output.
That is a self-report survey, so it describes what people say they do, and a person who thinks critically may also be a person who feels confident. It still points in the same direction as the usage logs. The behaviour that separates the two groups in both datasets is the work done after the model replies.
What beginners are measurably bad at#
The beginner side has a controlled study. Zamfirescu-Pereira, Wong, Hartmann and Yang (CHI 2023) gave people without AI expertise a purpose-built tool for designing prompts and watched them. Participants worked opportunistically rather than systematically, generalised from a single success or a single failure, and struggled to form an accurate model of how the system behaved. The study used 2023 models, and providers have since worked to make this difficulty smaller, so it shows what beginners did then and does not show that they still do it.
What nobody has tested#
Three gaps sit under every confident answer to this question. Nobody has measured expertise in a domain separately from time spent with the tool, so a surgeon new to chat models and a hobbyist with two years of daily use are not distinguishable in the logs. Nobody has compared the quality of work, as judged by someone other than the model, between iterative and directive users. And the one large field experiment on professional background points somewhere awkward: in the Procter and Gamble study by Dell'Acqua and colleagues, people using AI produced balanced proposals whatever their training, so the distinctive contribution of a technical or a commercial specialist was no longer visible in the output.
An inference: the expert's edge is the check they can run#
What follows is the SuperSkills reading, put forward as an inference and not as a result. Dell'Acqua and colleagues found in 2023 that consultants given a task just outside GPT-4's competence did worse with the tool than without it, and that the frontier between what the model does well and badly is local to a domain and has to be learned. Someone who knows a field can run that check from inside it. They notice that a figure is the wrong order of magnitude, that a clause does not fit the jurisdiction, that a plausible step skips the hard one. A beginner has no such position to check from, and the confident fluency of the output offers none.
If that reading is right, the Lee finding and the Anthropic pattern are two views of one habit. Self-trust gives a person a basis for pushing back, and pushing back is what iteration looks like in a log. It would also explain why the advice to draft first, discussed in human at the start, matters most for the people least able to spot a divergence. The reading is testable: vary a person's domain knowledge, hold the tool fixed, and measure who catches the planted error.
Borrowing the habit without the years#
The evidence supports three small practices, each of which costs minutes and none of which needs a model upgrade. Reply to the first output rather than accepting it, because the one dataset that exists says reply-and-revise is what the longer-serving users do. Form your own view before reading the answer, so that there is something to compare against. And keep a short list of the errors you have caught in your own field, since the frontier is local and a list of past catches is the nearest thing to a map of it. Whether these produce better work has not been measured. They are cheap, and the survey and the logs agree on where the difference lies.
For the case of someone early in a career, the same question looks different, because the person has not built the domain knowledge the check relies on. That is taken up at should juniors use AI at all and, as a design question about learning, at the expertise reversal effect.
Key sources
- Massenkoff, M., Lyubich, E., McCrory, P., Appel, R. and Heller, R. (2026). Anthropic Economic Index report: Learning curves. Anthropic, 24 March 2026. Read at source 2 October 2026; the headline and the three regression estimates are quoted from the report's own text. Graded entry.
- Lee, H.-P., Sarkar, A., Tankelevitch, L., Drosos, I., Rintel, S., Banks, R. and Wilson, N. (2025). The Impact of Generative AI on Critical Thinking. Microsoft Research and Carnegie Mellon, CHI 2025. Graded entry.
- Zamfirescu-Pereira, J. D., Wong, R. Y., Hartmann, B. and Yang, Q. (2023). Why Johnny Can't Prompt. CHI 2023. Graded entry.
- Dell'Acqua, F. et al. (2023). the BCG jagged-frontier field experiment. Harvard Business School and BCG working paper, 758 consultants; full title in the graded entry. Graded entry.
- Dell'Acqua, F. et al. (2025). The Cybernetic Teammate. NBER Working Paper 33641. Graded entry.
Related SuperSkills research#
On where the model's competence stops, the jagged frontier and how do I know when AI is wrong. On the habit of challenging the output, how do I get AI to challenge me and AI and critical thinking. On whether prompting skill is the thing to learn, why learning to prompt is weak career advice. On using the tool without wearing down your own capability, using AI without dependency.
Evidence review · SS-2026-385 · Graded against the published rubric · 1 vendor study
Hirji, R. (2026). How do expert AI users work differently from beginners?. The SuperSkills evidence base, SS-2026-385. https://thesuperskills.com/research/how-do-expert-ai-users-work-differently. Last reviewed 2 October 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work