Two large measurements, in two different countries, on two different kinds of homework, find the same shape of result. Giving a student an AI tool with no guardrails raises the mark on the homework itself and lowers what the student can do without it soon afterwards. The damage is neither universal nor even: it concentrates in whichever students actually hand the task over rather than working through it, and it can be cut roughly in half by how the tool is built, without taking the tool away.
Definition#
Homework outsourcing: using an AI tool to produce a homework answer rather than to help produce it, identified in the data by unusually short completion time paired with a normal or improved mark. It is a behaviour pattern found in the data, not a label applied to any particular student.
Turkey: grades up while the tool is there, down once it is removed#
Bastani and five colleagues ran a field experiment with nearly a thousand high-school students at a large school in Turkey across three arms: no AI, unrestricted GPT-4, and a version of GPT-4 built to tutor rather than to answer, withholding the full solution and drawing on a bank of common student mistakes. While the tools were available, unrestricted access raised grades by 48 per cent and the tutoring version by 127 per cent. Then the tools were taken away and the students sat an unassisted exam. The unrestricted group scored 17 per cent lower than the students who had never had AI access at all. The tutoring group kept most of what they had gained.
Same subject, same age group, same model underneath. The difference was entirely in what the interface let the student do with it.
China: thirty months, 26,811 students, and where the loss lands#
Strömberg, Lei and Wu tracked 26,811 Chinese secondary students across thirty months, combining homework scores and completion time with closed-book, invigilated exams in nine subjects. Homework scores rose 18 per cent after AI adoption and the time spent on it fell 30 per cent. Monthly exam scores, taken without the tool, fell 20 per cent within six months, and scores on the high-stakes entrance exams that decide secondary and university placement fell 18 to 24 per cent, with the full effect only visible after about two years.
The load-bearing detail is who carried the loss. It concentrated in the roughly 80 per cent of AI users whose pattern looked like outsourcing: homework finished unusually fast while scoring well. Students who kept working at close to their normal pace, even with the tool available, were largely spared. Losses were largest in social science, then STEM, then languages, and fell hardest on younger students, higher-achieving students and boys. This is a working paper on one country's school system, not a randomised trial, so the adoption pattern is self-selected rather than assigned; the size and duration of the panel is what makes it worth reading despite that.
The two studies agree on the mechanism and disagree on nothing that matters#
Different countries, different subjects, different methods: an experiment that could assign students to conditions, and a panel that could only watch what students already chose to do. Both land on the same account. Access to an AI tool is not, by itself, the thing that damages learning. What a student does with the access is. The Turkish experiment shows the mechanism at the level of one interface decision. The Chinese panel shows it at the scale of an entire school system tracked for two and a half years. Neither would be as convincing alone.
What a parent can actually do with this#
Completion time is the signal both studies point to, and a parent can watch for it without reading either paper. Homework finished much faster than usual, especially alongside marks that have not dropped, is the pattern both measurements associate with outsourcing rather than help. It is not proof on its own, in either study, but it is the same variable that separated the students who lost ground from the students who did not.
The Turkish result offers a second, more constructive lever. A tutoring-style tool that withholds the final answer and works through the method preserved almost all the benefit and avoided almost all of the measured harm. That is a choice about which AI product a school or a family uses, not a choice about whether to allow AI at all. Of everything either study measured, it is the one variable a parent or a teacher can actually set.
Key sources
- Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö. and Mariman, R. (2025). Generative AI Without Guardrails Can Harm Learning: Evidence from High School Mathematics. Proceedings of the National Academy of Sciences, 122(26).
- Strömberg, D., Lei, V. and Wu, Y. (2026). The Generative AI Learning Penalty: Evidence from Chinese Secondary Education. CEPR Discussion Paper 21577.
Related SuperSkills research#
On what the same mechanism does further along a career, how juniors become senior if AI does the junior work. On the family of designs that isolate delegation from access more generally, does using AI stop you learning. On the age-specific question this page does not answer, should children use AI. On the practice underneath both results, productive struggle.
About this research#
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. Both studies credited above are independent academic and working-paper research; nothing on this page is a SuperSkills coinage. The graded entries, with their stated limits, are in the evidence base.
Evidence review · SS-2026-361 · Graded against the published rubric
Hirji, R. (2026). What happens to homework when AI can do it?. The SuperSkills evidence base, SS-2026-361. https://thesuperskills.com/research/what-happens-to-homework. Last reviewed 27 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work