The AI Memo Test is a three-question self-assessment that Rahim Hirji published on 11 May 2025, a month after four chief executives sent their staff memos about artificial intelligence within days of each other. It scores what an organisation has decided. Seventeen months later the four firms whose memos prompted it have gone in different directions, and that divergence turns out to be the most useful thing on this page.
The answer, in one line
The AI Memo Test is a three-question self-assessment of whether an organisation has decided how AI changes its work.
Definition#
The AI Memo Test: a three-question self-assessment of whether an organisation has decided how AI changes its work, scored on tool fluency as a requirement, job descriptions written around what software already does, and managers held accountable for system-led productivity.
Three questions and four bands#
The questions are put to the organisation and answered yes or no.
- Have you made tool fluency a basic requirement across roles?
- Are your job descriptions written to avoid tasks software can already handle?
- Are your managers accountable for system-led productivity, not just team output?
Three out of three is ahead of the curve. Two is holding on. One is exposed. Zero is already behind. Hirji also maps the test onto a two-by-two of tool fluency against systems mindset: low on both is vulnerable, high fluency with no systems thinking is technical but limited, a strong mindset with low fluency is intuitive but unstructured, and strength in both is future-ready.
The test takes about a minute to run, which is its point. It was written for a leader who has read four memos in a month and wants to know whether the same conversation has happened in their own building.
The five words the Shopify memo does not contain#
The essay that introduces the test opens on Shopify, and describes Tobi Lütke emailing his staff five words: "AI is the new default." That sentence does not appear in the memo. Lütke published the memo himself on 7 April 2025, after it began to leak, and it runs to roughly 1,300 words. Its own title and its load-bearing line read: "Reflexive AI usage is now a baseline expectation at Shopify." The memo also told staff to demonstrate why AI could not do a job before asking for more headcount, and said AI use would enter performance and peer reviews.
The five-word version is a fair compression of what the memo argues and it carries quotation marks around words the memo never used. This page therefore quotes the memo and not the compression. The correction is recorded here because the estate checks other people's figures and has to check its own; the same discipline produced two corrections to Hirji's own essay on the missed reps.
Four similar memos, four different outcomes#
The four firms wrote in the same register within a month. What followed separates them.
Shopify. The position stood. The memo remains the clearest published statement of the first question, since it makes AI fluency a condition of resourcing and a line in a performance review. Graded entry.
Duolingo. Luis von Ahn announced an AI-first approach in late April 2025 and said the company would gradually stop using contractors for work AI can handle. On 23 May 2025 he published a clarification, opening with the observation that one of the most important things leaders can do is provide clarity and that he had not done it well. He added that he did not see AI as replacing what employees do, and that hiring was continuing at the same speed as before. He said later that Duolingo never laid off any full-time employees. Graded entry.
Fiverr. Micha Kaufman told staff on 7 April 2025 that AI was coming for their jobs and for his own, and that those who did not adapt would face a career change within months. On 17 September 2025 he announced that the company would part with approximately 250 team members, about 30 per cent of its workforce, describing a painful reset. Graded entry.
Box. Aaron Levie's memo was posted as an image on X in early May 2025. It could not be opened and read at source for this page, so it is described here from the essay that cites it and carries no graded entry. Levie's stated position since has been that agents are not taking anybody's job at Box.
Four documents written in the same month, in the same voice, about the same technology, preceded a walk-back, a redundancy of roughly a third, a position that held, and a position that softened. The memos did not predict which was which.
The test scores a decision, and a decision is not a result#
Every one of the three questions asks what an organisation has settled. None asks what happened afterwards. That is a reasonable design for a one-minute instrument and it sets a hard limit on what a score means. A firm scoring three out of three has decided three things. Whether those decisions produced capability, or produced the appearance of it, sits outside the test entirely.
The estate has a name for the gap. Usage theatre is activity that satisfies a policy without changing the quality of the work underneath, and question three invites it directly: a manager held accountable for system-led productivity, with no accompanying measure of whether the judgement in the work survived, will optimise for the number that is counted. The METR result is the standing warning here, since the developers in it were slower with the tool and believed they had been faster.
Question two has a cost the test does not price#
"Written to avoid tasks software can already handle" is a sound instruction for a job description and it is the same instruction as "remove the tasks juniors learn on". The tasks a model handles well are disproportionately the small, specifiable, reviewable ones, and those are the reps by which a person becomes senior. Bastani and colleagues found that students given unrestricted access performed 17 per cent worse than a control once the tool was taken away, while a guardrailed version of the same tool largely removed the harm. The design of the interface decided the outcome, and the same logic applies to the design of a role. Graded entry.
So a firm can answer yes to question two and be building capability debt, by writing roles that deliver this quarter and produce nobody able to supervise the work in five years. The estate's treatment of that sits at the missing rungs and how juniors become senior. Scoring the question yes is the beginning of the question, not the end of it.
Nothing validates the three questions#
No study tests whether the three questions measure one thing, whether the four bands separate anybody, or whether a score predicts any outcome at all. The two-by-two has never been applied to a sample. The divergence of the four firms above does not supply the missing test either: four firms, no control group, and outcomes confounded by market conditions, sector and the ordinary business of running a company.
That places the AI Memo Test in the same category as the five-question self-score published alongside it. Both are prompts for a conversation with a number attached. Neither is an index, and an instrument that has never been tested should say so on its face rather than in a footnote.
Running it so the score earns its keep#
Score the three questions honestly, then put a second question against each yes.
- Fluency as a requirement. Required of whom, assessed how, and what happens to somebody who does not meet it? A requirement nobody is measured against is a value statement.
- Job descriptions rewritten. Which tasks came out, and where does a person now acquire the judgement those tasks used to teach? If the answer is nowhere, the description has moved a cost into the future.
- Managers accountable for system-led productivity. Accountable against which measure, and what sits beside it to catch a drop in quality? Without the second measure, question three rewards the appearance of adoption.
Then read the four firms again. Every one of them would have scored well in May 2025, and their paths since had more to do with what they built underneath the memo than with the memo itself.
Key sources
- Hirji, R. (2025). The AI Memo Test. Box of Amazing, 11 May 2025. The dated first publication of the term.
- Lütke, T. (2025). Shopify memo on AI usage, published by the author on X, 7 April 2025; text as reported. Graded entry.
- von Ahn, L. (2025). AI-first memo and the clarification of 23 May 2025. Graded entry.
- Kaufman, M. (2025). Fiverr memo of 7 April 2025 and the workforce reduction of September 2025. Graded entry.
- Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö. and Mariman, R. (2025). Generative AI without guardrails can harm learning. PNAS, 122(26). Graded entry.
Related SuperSkills research#
The companion instrument for an individual is how good are you at using AI. On what a high adoption score can conceal, usage theatre and how to measure AI adoption properly. On the cost inside question two, the missing rungs, capability debt and how juniors become senior. The wider argument about deciding rather than drifting sits at drift versus design.
About this research#
Written by Rahim Hirji, author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. The AI Memo Test is his, with a dated first publication of 11 May 2025 read at source. This page corrects one attribution in that essay: the five words given as a quotation from the Shopify memo do not appear in it, and the memo's own sentence is used instead. Findings are attributed to the studies that produced them and kept separate from the interpretation.
Explainer · SS-2026-223 · Graded against the published rubric
Hirji, R. (2026). What is the AI Memo Test?. The SuperSkills evidence base, SS-2026-223. https://thesuperskills.com/research/what-is-the-ai-memo-test. Last reviewed 12 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work