A hiring process is a measuring instrument, and most of what it measures is effort. The covering letter was never read for its prose. It was read as evidence that somebody had spent an evening on this employer rather than on forty others. The work sample was read the same way: not as an artefact but as a trace of the hours behind it.
When the cost of producing a good trace falls to nothing, the instrument stops separating candidates, and it does so quietly. Nothing breaks. Applications go up. The forms are better written than they have ever been. What has gone is the information that made them worth reading, and no report inside a hiring function measures that.
The firm's response makes the problem symmetrical#
The reflex is to screen with the same technology, at the same speed, on the same volume. That restores throughput. It does not restore information, because both sides are now using a tool that raises the floor, and a floor that rises on both sides of a transaction leaves the ranking roughly where it was while costing both parties more to produce.
New York City has already written the specific failure into binding rules. Local Law 144 became law in December 2021 and came into force in January 2023; the implementing rules made under it, at 6 RCNY Subchapter T, define the trigger for an automated employment decision tool in three limbs. The third is to use a simplified output to overrule conclusions derived from other factors including human decision-making. A city agency drafting rules about hiring software described the mechanism by which a machine score displaces a human judgement about a person, and did so before most employers had a policy on it.
What happened next matters more than the rule. The New York State Comptroller audited the first two years of enforcement and found that the city department surveyed 32 companies and identified a single issue of non-compliance, while the auditors reviewing the same companies identified at least seventeen instances of potential non-compliance. Two complaints were received in two years. The audit says potential non-compliance rather than violation, and the department disputes elements of the finding. A rule existing and a rule operating are separate facts, and only one of them was ever in doubt.
Why output is a weaker measure than it was#
Three findings in the graded base bear on this directly, and none of them is about hiring.
The floor rises faster than the ceiling. Brynjolfsson, Li and Raymond followed 5,172 customer-support agents through a staggered rollout and found resolutions per hour up about 15 per cent on average, 30 per cent for the less skilled and less experienced, rising to 36 per cent in the lowest skill quintile, with no significant gain for the most skilled. A tool with that shape compresses the visible distance between a strong and a weak performer at exactly the point in the distribution where entry-level selection happens. The study measures output in one firm and one occupation over months, and says nothing about whether those novices went on to become experts.
A work sample measures the task as much as the person. Dell'Acqua and colleagues gave 758 consultants tasks inside and just outside the model's competence. Inside, assisted consultants were markedly better and faster. Outside, they did worse than consultants working with no AI at all. Set a candidate a task that sits inside the frontier and the exercise tells you where the task sits. Set one that straddles it and the exercise starts telling you about the candidate again.
The rungs are being removed through hiring rather than through dismissal. Hosseini Maasoum and Lichtinger, working across 281,111 firms and roughly 66 million workers, find junior employment at adopting firms about 9 per cent below non-adopters six quarters after diffusion, with no comparable break in senior employment, concentrated in the exposed occupations. The authors are careful throughout: they describe adoption as associated with the decline and call the evidence suggestive rather than causal. Read with that qualifier intact, it still describes a labour market where the entry route narrows without anybody being let go, and an unemployment statistic cannot see a job that was never advertised.
What this asks of a hiring process#
- Decide what the exercise is for before choosing it. An exercise that tests production now tests access to a tool. An exercise that tests judgement about production still tests the person.
- Make at least one stage unaided, and say so in advance. Candidates who are told will prepare for it, which is the behaviour being selected for.
- Ask the candidate to account for their own work. The question that separates a capable candidate from a well-tooled one is not what they produced but which decisions they made inside it, and where they disagreed with what came back.
- Treat the volume as a cost, not a compliment. A tenfold rise in applications with no rise in hires is a tax on the hiring team, paid in attention.
- Write down what you screen with. If a tool contributes to a rejection, somebody will eventually ask which tool, on what basis, audited by whom. New York has shown that the asking can take years and arrive from a direction nobody was watching.
What would settle it#
Application volume against subsequent performance, measured inside firms, before and after the tools became general. Employers hold both halves of that dataset and nobody publishes it. Until somebody does, the claim on this page is an argument about measurement rather than a finding about hiring, and the three studies above are evidence about adjacent things.
Where this sits in my own argument#
This is the front-door version of synthetic seniority: output that looks senior while the judgement underneath was never built. Assessment at the point of entry is the first place a firm can notice, and the place most firms have changed least. The structural half is the missing rungs, and the choice between fixing the instrument and leaving it is the ordinary form of drift versus design.
Related SuperSkills research#
On measuring the person rather than the artefact, how do you assess capability rather than output? On the machine at the front of the funnel, should AI reject an application before a human reads it? On the entry-level evidence, will AI replace entry-level jobs? and should juniors use AI?
Key sources
- Brynjolfsson, E., Li, D. and Raymond, L. (2023). Generative AI at Work.
- Dell'Acqua, F. et al. (2023). Field experimental evidence on the jagged technological frontier.
- Hosseini Maasoum, S. M. and Lichtinger, G. (2026). Generative AI as Seniority-Biased Technological Change.
- Council of the City of New York (2021). Local Law 144, automated employment decision tools.
- Office of the New York State Comptroller (2025). Enforcement of Local Law 144.
About this research#
Written by Rahim Hirji, author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company.
How this research works · Reviewed quarterly · Found an error? Tell me and it is corrected on the page.
Evidence review · SS-2026-389 · Graded against the published rubric · 2 working papers, 1 peer-reviewed study, 1 statutory investigation and 1 of other kinds
Hirji, R. (2026). Hiring Signals When Effort Is Free. The SuperSkills evidence base, SS-2026-389. https://thesuperskills.com/research/hiring-signals-when-effort-is-free. Last reviewed 3 October 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work