← Research
Research

What is the AI Memo Test?

Three questions, four bands, and four firms that no longer agree with each other.

Last reviewed: 12 September 2026

The three questions and where the score lands, the five words attributed to the Shopify memo that the memo does not contain, what Shopify, Duolingo, Box and Fiverr each did in the seventeen months afterwards, and the reason a score of three out of three is a statement about a decision and not about an outcome.

Question this page answersAll 811 questions this research covers

The AI Memo Test is a three-question self-assessment that Rahim Hirji published on 11 May 2025, a month after four chief executives sent their staff memos about artificial intelligence within days of each other. It scores what an organisation has decided. Seventeen months later the four firms whose memos prompted it have gone in different directions, and that divergence turns out to be the most useful thing on this page.

The answer, in one line

The AI Memo Test is a three-question self-assessment of whether an organisation has decided how AI changes its work.

Share as a card

Definition#

The AI Memo Test: a three-question self-assessment of whether an organisation has decided how AI changes its work, scored on tool fluency as a requirement, job descriptions written around what software already does, and managers held accountable for system-led productivity.

Share this definition as a card

Three questions and four bands#

The questions are put to the organisation and answered yes or no.

  1. Have you made tool fluency a basic requirement across roles?
  2. Are your job descriptions written to avoid tasks software can already handle?
  3. Are your managers accountable for system-led productivity, not just team output?

Three out of three is ahead of the curve. Two is holding on. One is exposed. Zero is already behind. Hirji also maps the test onto a two-by-two of tool fluency against systems mindset: low on both is vulnerable, high fluency with no systems thinking is technical but limited, a strong mindset with low fluency is intuitive but unstructured, and strength in both is future-ready.

The test takes about a minute to run, which is its point. It was written for a leader who has read four memos in a month and wants to know whether the same conversation has happened in their own building.

The five words the Shopify memo does not contain#

The essay that introduces the test opens on Shopify, and describes Tobi Lütke emailing his staff five words: "AI is the new default." That sentence does not appear in the memo. Lütke published the memo himself on 7 April 2025, after it began to leak, and it runs to roughly 1,300 words. Its own title and its load-bearing line read: "Reflexive AI usage is now a baseline expectation at Shopify." The memo also told staff to demonstrate why AI could not do a job before asking for more headcount, and said AI use would enter performance and peer reviews.

The five-word version is a fair compression of what the memo argues and it carries quotation marks around words the memo never used. This page therefore quotes the memo and not the compression. The correction is recorded here because the estate checks other people's figures and has to check its own; the same discipline produced two corrections to Hirji's own essay on the missed reps.

Four similar memos, four different outcomes#

The four firms wrote in the same register within a month. What followed separates them.

Shopify. The position stood. The memo remains the clearest published statement of the first question, since it makes AI fluency a condition of resourcing and a line in a performance review. Graded entry.

Duolingo. Luis von Ahn announced an AI-first approach in late April 2025 and said the company would gradually stop using contractors for work AI can handle. On 23 May 2025 he published a clarification, opening with the observation that one of the most important things leaders can do is provide clarity and that he had not done it well. He added that he did not see AI as replacing what employees do, and that hiring was continuing at the same speed as before. He said later that Duolingo never laid off any full-time employees. Graded entry.

Fiverr. Micha Kaufman told staff on 7 April 2025 that AI was coming for their jobs and for his own, and that those who did not adapt would face a career change within months. On 17 September 2025 he announced that the company would part with approximately 250 team members, about 30 per cent of its workforce, describing a painful reset. Graded entry.

Box. Aaron Levie's memo was posted as an image on X in early May 2025. It could not be opened and read at source for this page, so it is described here from the essay that cites it and carries no graded entry. Levie's stated position since has been that agents are not taking anybody's job at Box.

Four documents written in the same month, in the same voice, about the same technology, preceded a walk-back, a redundancy of roughly a third, a position that held, and a position that softened. The memos did not predict which was which.

The test scores a decision, and a decision is not a result#

Every one of the three questions asks what an organisation has settled. None asks what happened afterwards. That is a reasonable design for a one-minute instrument and it sets a hard limit on what a score means. A firm scoring three out of three has decided three things. Whether those decisions produced capability, or produced the appearance of it, sits outside the test entirely.

The estate has a name for the gap. Usage theatre is activity that satisfies a policy without changing the quality of the work underneath, and question three invites it directly: a manager held accountable for system-led productivity, with no accompanying measure of whether the judgement in the work survived, will optimise for the number that is counted. The METR result is the standing warning here, since the developers in it were slower with the tool and believed they had been faster.

Question two has a cost the test does not price#

"Written to avoid tasks software can already handle" is a sound instruction for a job description and it is the same instruction as "remove the tasks juniors learn on". The tasks a model handles well are disproportionately the small, specifiable, reviewable ones, and those are the reps by which a person becomes senior. Bastani and colleagues found that students given unrestricted access performed 17 per cent worse than a control once the tool was taken away, while a guardrailed version of the same tool largely removed the harm. The design of the interface decided the outcome, and the same logic applies to the design of a role. Graded entry.

So a firm can answer yes to question two and be building capability debt, by writing roles that deliver this quarter and produce nobody able to supervise the work in five years. The estate's treatment of that sits at the missing rungs and how juniors become senior. Scoring the question yes is the beginning of the question, not the end of it.

Nothing validates the three questions#

No study tests whether the three questions measure one thing, whether the four bands separate anybody, or whether a score predicts any outcome at all. The two-by-two has never been applied to a sample. The divergence of the four firms above does not supply the missing test either: four firms, no control group, and outcomes confounded by market conditions, sector and the ordinary business of running a company.

That places the AI Memo Test in the same category as the five-question self-score published alongside it. Both are prompts for a conversation with a number attached. Neither is an index, and an instrument that has never been tested should say so on its face rather than in a footnote.

Running it so the score earns its keep#

Score the three questions honestly, then put a second question against each yes.

Then read the four firms again. Every one of them would have scored well in May 2025, and their paths since had more to do with what they built underneath the memo than with the memo itself.

Key sources

The companion instrument for an individual is how good are you at using AI. On what a high adoption score can conceal, usage theatre and how to measure AI adoption properly. On the cost inside question two, the missing rungs, capability debt and how juniors become senior. The wider argument about deciding rather than drifting sits at drift versus design.

About this research#

Written by Rahim Hirji, author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. The AI Memo Test is his, with a dated first publication of 11 May 2025 read at source. This page corrects one attribution in that essay: the five words given as a quotation from the Shopify memo do not appear in it, and the memo's own sentence is used instead. Findings are attributed to the studies that produced them and kept separate from the interpretation.

Explainer · SS-2026-223 · Graded against the published rubric

Cite this page

Hirji, R. (2026). What is the AI Memo Test?. The SuperSkills evidence base, SS-2026-223. https://thesuperskills.com/research/what-is-the-ai-memo-test. Last reviewed 12 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

What is the AI Memo Test?

The AI Memo Test is a three-question self-assessment of whether an organisation has decided how AI changes its work. The questions are whether tool fluency is a basic requirement across roles, whether job descriptions are written to avoid tasks software already handles, and whether managers are accountable for system-led productivity and not only for the output of their team. Rahim Hirji published it in Box of Amazing on 11 May 2025.

What are the four bands in the AI Memo Test?

Three out of three is ahead of the curve, two is holding on, one is exposed, and none is already behind. The bands describe how many of the three decisions an organisation has taken. They carry no validation: nothing tests whether a score of three predicts a better outcome than a score of one.

Which memos prompted the AI Memo Test?

Four sent within a month of each other in spring 2025. Tobi Lutke at Shopify on 7 April, whose published memo states that reflexive AI usage is now a baseline expectation at the company. Micha Kaufman at Fiverr, also on 7 April, who told staff that AI was coming for their jobs and his own. Luis von Ahn at Duolingo in late April, announcing an AI-first approach and a gradual end to contractors doing work AI can handle. And Aaron Levie at Box, posted as an image on X in early May.

Did the Shopify memo say AI is the new default?

No. The original essay introducing the AI Memo Test describes Lutke emailing his staff the five words AI is the new default. The memo Lutke published on X on 7 April 2025 runs to roughly 1,300 words and its load-bearing sentence reads: Reflexive AI usage is now a baseline expectation at Shopify. The five-word version is a fair compression of what the memo argues and it is not a quotation from it, so this page uses the memo's own wording.

What happened to the four companies afterwards?

They diverged. Shopify's position stood. Von Ahn published a clarification on 23 May 2025 saying he had not provided clarity and that he did not see AI as replacing what employees do, with hiring continuing at the same speed, and said later that Duolingo never laid off any full-time employees. Kaufman announced in September 2025 that Fiverr would part with about 250 people, roughly 30 per cent of its workforce. Levie's stated position at Box has remained that agents are not replacing anybody's job. Four memos with similar wording produced four different outcomes, which is the strongest available caution against reading a memo as a forecast.

In this hub

The named concepts

The vocabulary this research contributed, and what each term does.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory for CEOs and boards  ·  Enquire