← Research
Research

Proving you did the work

Every process that certified a person now certifies an artefact anybody can produce. Detection fails, and it fails hardest on the people who can least afford it.

Last reviewed: 1 September 2026

Liang and colleagues found seven widely used GPT detectors misclassified over half of non-native English essays while classifying US schoolchildren almost perfectly. Yin and colleagues found AI replies made people feel more heard than untrained humans until the reply was labelled. What survives is proof of process rather than proof of absence.

Questions this page answersAll 529 questions this research covers

Every process that used to certify a person now certifies an artefact anybody can produce. The application letter, the take-home exercise, the portfolio, the written submission: each was a proxy for capability, and each has become cheap. The instinct is to detect the machine. That instinct fails, and it fails hardest on the people who can least afford it.

Detection fails, and it fails unevenly#

Liang and colleagues evaluated seven widely used GPT detectors against TOEFL essays written by non-native English speakers and essays written by US eighth-grade students. The detectors classified the American schoolchildren with near-perfect accuracy. They misclassified more than half of the non-native essays as AI-generated, an average false positive rate of 61.22 per cent.

The proposed mechanism explains why this is not a tuning problem. Detectors lean on perplexity, a measure of how predictable text is, and second-language writing is more predictable: smaller vocabulary, more conventional constructions, fewer idiosyncratic turns. The property being measured as machine-likeness is a property of writing carefully in a language you learned later. Vendors dispute how far this generalises to current tools, and the study tested what was available at the time. The mechanism has not gone away.

So an organisation running detection is not running a neutral check. It is running one that penalises second-language speakers at scale, and doing so invisibly, because a false accusation looks identical to a true one from the outside. MIT’s own committee reached the same practical conclusion from a different direction, recommending against AI detectors on the grounds that it invites an arms race.

The measurable cost of disclosure#

The obvious alternative to detection is disclosure. Before adopting it as policy, it is worth knowing what disclosure does.

Yin, Jia and Wakslak compared AI-generated and human-written replies, with and without disclosure. AI-generated replies made recipients feel more heard than replies from untrained humans. Labelling the reply as AI removed that advantage entirely.

Read the shape of that carefully, because it is not an argument for concealment. The same words produced a different effect depending on what the recipient believed about their origin. What people value in being responded to has little to do with the quality of the sentence. They value the belief that somebody chose to attend to them. Disclosure removes that belief and the benefit goes with it, even where the disclosed work was better. The authors also note that norms here are moving and nobody has measured the effect over time, so treat the size of the penalty as unstable rather than fixed.

Provenance is arriving, built for somebody else#

From 2 August 2026, Article 50(2) of the EU AI Act requires providers of systems generating synthetic text, audio, image or video to mark outputs in a machine-readable format so they are detectable as artificially generated. The obligation sits on the model provider. Education, employment and hiring are not mentioned anywhere in it.

Two consequences follow for anybody trying to prove their own work. Marking will establish, increasingly reliably, that a passage was generated. It will never establish that the person submitting it understood a word of it. And Article 50(2) exempts systems performing an assistive function for standard editing or not substantially altering the input, covering most legitimate professional use. The regime is being built to answer a question about content, not a question about a person.

What actually works: proof of process, not proof of absence#

Every approach above tries to prove a negative, that a machine was not involved. That is unprovable, increasingly so, and biased in its failures. The approaches that survive contact with reality all share a shape: they demonstrate that the person can operate the knowledge, rather than that they generated the artefact.

Hiring when everybody uses AI#

The application letter is finished as a signal and pretending otherwise wastes everybody’s time. What replaces it is not a better filter but a different one.

The practical move is to stop screening on produced artefacts and start screening on live reasoning, earlier in the process than is comfortable. That costs more per candidate, which is the fair objection to it. It is also the only approach that does not disadvantage the second-language applicant twice: once through detection, and once through a written filter that rewards fluency over competence. If you are going to spend, spend on a short conversation rather than a longer take-home exercise.

On putting AI skills on a CV, the position from why learn to prompt is weak career advice holds: the skill is neither scarce nor durable, so it reads as a claim about tooling rather than about capability. State instead a specific thing you built or decided with these systems, and what you were accountable for when it went wrong. That is a claim about judgement, and judgement is the part that does not commoditise.

Should employees disclose?#

Yes, but the policy has to be specific about what, and most are not. A blanket requirement to declare any AI involvement is unenforceable, produces meaningless declarations on almost everything, and imposes the Yin penalty on work whose quality was never in question.

The defensible version discloses by function rather than by tool. Declare it where somebody is entitled to know a human attended personally: condolences, apologies, feedback on a person’s work, anything where the point of the message is that a human chose to send it. Declare it where accountability transfers, meaning any output somebody else will rely on without checking. Do not require it for drafting, research or summarising, where the tool is doing what a search engine or a template did before, and where the declaration carries a real social cost for no informational gain.

Where this sits in my own argument#

The reason I keep returning to verification is that it is where the cost lands and where nobody is looking. Checking somebody else’s output is harder than producing your own, pays less, and is almost never in a job description. That is the argument in verification work, and this page is what happens when the same problem reaches a hiring process.

My position is that every attempt to prove a machine was absent will fail, and that the effort should move to demonstrating that a person can operate the knowledge. That is also the honest test for synthetic seniority: output that looks senior while the judgement underneath was never built.

What I have observed in organisations#

Proving you did the work has become genuinely difficult, and the mistrust runs in every direction. Anything that reads as machine-written attracts suspicion, from a senior colleague, from a junior, and from clients. An em dash is now treated as evidence.

The damaging part is the generalisation. Once somebody has seen one piece of obviously generated work, they start reading everything that way, including work that was done properly by a person who happens to write cleanly.

Meanwhile teams are swirling around clearing up after each other. A great deal of current effort goes into fixing generated material that arrived looking finished. Anecdotally, and I offer it as no more than that, some of this work now takes about as long as it did before, once the clean-up is counted. That is uncomfortably close to what METR measured above, and the reason it goes unnoticed is the same: nobody is counting the second pass. What is missing in most of these teams is not a tool. It is clarity about how the work is supposed to be done.

What this page does not answer#

Who owns AI-generated work is a legal question rather than an evidential one, it varies by jurisdiction and by how the output was produced, and this estate holds no graded source on it. It is left open rather than answered from general knowledge, because the wrong answer here is expensive and confidently given everywhere.

Nor does this page claim the four process methods above are proven. They are the ones consistent with the evidence on what detection and disclosure actually do, and with the mechanism that makes explanation harder to fake than production. Nobody has run the trial that compares them.

Cite this

Hirji, R. (2026). Proving you did the work. The SuperSkills Intelligence Company. Last reviewed 1 September 2026. thesuperskills.com/research/proving-you-did-the-work

In this hub

Work, careers and the labour market

What happens to jobs, careers and the first rung.

The work

Where the writing comes from.

These essays draw on research across more than 200 organisations in 30 countries. See the wider body of work, or bring it into your organisation.

All research →
Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory and coaching  ·  Enquire