Short version
A 30-month study of 26,811 students in central China found that when they started using AI for homework, their homework scores rose 18% and their homework time fell about 30% — and then their closed-book exam scores fell 20% within six months, and the entrance exams that decide their school and their university fell 24% and 18% (paper abstract, via Stanford SCALE). The strongest students lost the most. The children who escaped were the ones who took no time saving at all. Weeks later, and without citing the study, New York City stopped student-facing generative AI for about 600,000 children from 2-K through eighth grade (NYC Mayor's Office).
The study in three lines
- The scale 26,811 students, 30 months, nine subjects, real entrance exams Source: South China Morning Post
- The finding Homework up 18%, exams down 20%, entrance exams down up to 24% Source: The Decoder
- The consequence New York: 600,000 children, 2-K to grade 8, one year, from 2026–27 Source: NYC Mayor's Office
What the study did, and what it found
The paper is called The Generative AI Learning Penalty: Evidence from Chinese Secondary Education. David Strömberg of Stockholm University wrote it with Victor Lei and Yanhui Wu of the University of Hong Kong, and it covers 26,811 students in grades 7 to 12 — ages roughly 12 to 18 — in a single county in central China, over 30 months of monthly data (abstract, via Stanford's SCALE repository). The South China Morning Post, which broke the story, dates the window September 2022 to June 2025, and reports that around 80% of the students took up generative AI at some point in it (SCMP). The researchers had four things for each child: the marks on their homework, how many minutes the homework took, their monthly closed-book tests, and — at the end — the national exams.
That last item is what makes this different from almost everything else written on the subject. The outcome is not a task the researchers invented. It is the exam the child sat.
Two words decide how much weight the numbers carry. It is a working paper — research released before peer review, on SSRN as abstract 6868618 and as CEPR discussion paper DP21577; no journal has signed off on it. And the method is difference-in-differences: it compares how a child's results moved after they started using AI against how results moved for children who had not yet started, which strips out anything that hit the whole county at once (PPC Land names the estimator — Callaway and Sant'Anna, errors clustered at class level). What it cannot do is randomise. Nobody assigned these children to use AI.
Two of those rows need context a European reader will not have. The zhongkao and the gaokao are not end-of-year tests. They are single sittings that decide which secondary school a Chinese child attends and which university they can enter, with no realistic retake for most families. For students who had been using AI for two years or more, the paper puts the zhongkao fall at 24% and the gaokao fall at 18% (PPC Land, quoting the paper's figures; the same split is reported by SCMP). A 24% fall there is not a mark on a report card. It is a different school, and often a different life.
Notice the order of events, because the whole argument sits in it. The improvement arrives immediately and it is real: better homework, nineteen minutes back every evening, a calmer house. The damage arrives later, in a place the family is not looking.
Why this study is worth more than the usual one
Most of what gets written about AI and learning rests on a few dozen volunteers doing an invented task for an afternoon. This is 26,811 people over 30 months (abstract), and the outcomes are the real exams those children sat, the ones that decided their futures. When the measure is something the subjects care that much about, nobody phones it in. The limitations: it is a working paper rather than peer-reviewed research, nobody was randomly assigned, and it covers one county. Take the mechanism as the finding and the exact percentages as local.
Claim one: the time saved and the learning lost are the same thing
There is a group in the data with no measurable damage, and it is usually reported as the hopeful part. Read it properly and it is the bleakest finding in the study.
Students who used AI and still spent about as long on their homework as the non-users — 50 to 65 minutes, the range non-users worked in — came out with exam scores the paper describes as effectively identical to the non-users': same medians, same spread, despite scoring markedly higher on the homework itself (PPC Land, on the paper's distribution figure). The paper's own abstract is a shade more cautious, calling their losses "small" rather than absent (abstract) — either way, this is the group the harm skipped.
Now notice what that group actually did: they captured no time saving whatsoever. The children who escaped the harm were the ones who gave up the thing everyone buys the tool for. So the saving and the loss are not two things that happen to travel together — they are the same thing seen from two sides. The nineteen minutes were not waste being trimmed. They were where the learning happened.
The fingerprint is in the homework logs: after more than five months of AI use, about 81% of students were finishing their homework in under 50 minutes — faster than even the quickest non-users (The Decoder). And the more they used it, the worse the exams got: about a 5% loss for students using AI up to an hour a week, about 30% for those using it five hours or more (same source). The tool did not make the material easier to understand. It made it faster to be finished with.
This is why "just teach them to use it responsibly" is not the easy answer it sounds like. Responsible use, as this study defines it, means your child works exactly as long as they did before and gets no time back. That is a real trade, worth making — but it is a trade, not a technique.
Claim two: it is the good students who are in danger, and you will not find out for two years
The comfortable version of this story is that AI props up the strugglers while the strong students were always going to be fine. The data says the opposite.
The top third of students by prior attainment lost around 24%. The bottom third lost 16% (The Decoder's breakdown, which also puts boys at −21.6% against girls at −18.4%). The plausible reading is that the strong students had more to lose: a working habit of pushing through difficulty that was producing their results, and it was precisely that habit the tool replaced. The weaker students were not doing that work in the first place, so there was less to displace.
The subject pattern says the same thing. The smallest loss was in the children's own language, at 9% — competence built continuously by living. The largest was social sciences, at 27%, with STEM at 22% and English at 17% in between (The Decoder). What social science teaches is the ability to construct an argument. That is the exact task an AI will complete on your behalf, cheerfully, in seconds.
And then the timing, which is what makes this genuinely difficult. Monthly test scores moved within six months. The exams that decide a child's future took about two years to show the full damage (The Decoder). Through all of it the visible signals improved: better homework marks, less time at the table, calmer evenings.
Which is the sentence to take away from the whole study. A good report card is not evidence that nothing is happening. For two years, it is exactly what the damage looks like.
Claim three: New York's ban does not cite this study, but it cuts along the same lines
A city moved weeks after the largest study on the question landed — but say plainly what the city did not say: neither the mayor's announcement nor Chalkbeat's report on it mentions the Chinese study at all. What follows is not a claim about cause. It is that the cuts New York made fall where this evidence says they should, which is a fair test of a policy either way.
New York City has put a one-year stop on student-facing generative AI for every child from 2-K through eighth grade — about 600,000 of them, two thirds of the system — starting with the 2026–27 school year, and barred companion chatbots at every grade, high school included. Teachers may still use AI for instructional planning and operational tasks, and there are exceptions for assistive technology, for multilingual learners, and for career-readiness programmes such as computer science (the city's announcement, 2 September 2026). Chalkbeat adds the half the release does not spell out: staff may plan, translate and draft with AI, but not grade, monitor behaviour, counsel students or write special-education plans (Chalkbeat). Mayor Mamdani was blunt about the reasoning: "The tech industry wants us to believe that A.I.-powered early education is not only inevitable, but necessary. We do not see it that way."
The student/teacher line is the right cut. The damage in the Chinese data did not come from AI being present in a classroom. It came from the tool absorbing the child's practice. A teacher generating a lesson plan does not touch that mechanism. A child generating the answer is that mechanism.
The age line matches too, which is more than most policies manage. In the study the younger students lost around 24% against 17% for the older ones (The Decoder); the paper's abstract puts it as losses being most pronounced among junior students (abstract). New York stopped at eighth grade — roughly where that curve says the harm is worst.
Two things are bundled in that this study says nothing about, and they should be separated when you argue about it. The screen-time limits — no one-to-one device use at second grade and below, 30 minutes a day in grades 3 to 5, and 45 minutes in grades 6 to 8 (NYC Mayor's Office) — are an older debate with their own evidence. The companion-chatbot ban across all grades is about emotional dependence, not learning. Both may be reasonable; neither follows from this data.
The honest objection is the expiry date. In twelve months somebody decides what happens next. Chancellor Samuels names the standard — "Innovation does not mean more technology, and over the next year, we will lead with evidence to make sure technology serves learning" (same release) — but not which evidence, gathered how. A pause that simply lapses will have bought a year and learned nothing.
What the coverage disagrees on
Three things are not settled. Anyone repeating these numbers in a meeting should know which of them can be pushed back on.
How many students. The paper says 26,811 (abstract). The SCMP rounds to "more than 26,000"; Al Jazeera rounds the other way, to "about 27,000". Use the paper's figure and the argument survives either rounding.
Whether the time-matched group was unharmed or only lightly harmed. The paper's abstract says those users "experience small learning losses". Its distribution figure, as read by PPC Land, shows their medians and interquartile ranges effectively identical to non-users', and The Decoder writes flatly that they "scored just as well". The honest version of Claim One is the smaller one: matching your old homework time removes most of the penalty, not provably all of it.
Whether the penalty is growing or shrinking. This one cuts against the story. For any given student the full entrance-exam damage takes about two years to appear — but measured across the county by calendar date, the estimated penalty fell from about 25% in early 2023 to 16% by June 2025 (The Decoder). Later adopters may simply have used the tools better. Either way, anyone quoting 24% as today's figure is quoting the worst year.
What to do about it on Monday
A school can draw a line at the gate. The part that decides how your child turns out happens at your table, and the study points at things that are actually within your control.
- Insist on the order, not abstinence. Their attempt first, however rough; the tool afterwards. Being stuck after a genuine try and asking for help is how anyone has ever learned. Starting with the answer is the thing that was measured.
- Ask them to explain it with the screen closed. Not as a test — as a normal question over dinner. Thin explanation, thin understanding. It takes ninety seconds and it is the only exam you have.
- Treat homework that suddenly got fast as information, not as a crime. Something changed. Find out what.
- Watch the strong student hardest. The counter-intuitive finding is the useful one. Doing well is not protection; in this data it was a risk factor.
- Take the subject list seriously. If your child is writing essays with it, that is the case the evidence is most worried about.
Bottom line
Twenty-six thousand children handed their homework to an AI and came out doing worse on the exams that decided where they went next. Every visible sign said it was working, right until they had to do it alone.
New York's answer is blunt, and for eight-year-olds blunt is probably right: safe use needs somebody to notice whether a child is doing the work or performing it, and no school system can do that six hundred thousand times. A parent, for one child, can.
The evidence does not say keep them away. It says the struggle is what you spend when you buy the nineteen minutes back — and that nobody will send you the bill for two years.
FAQ
What did the study actually find?
Students who adopted AI scored 18% higher on their homework and finished it about 30% faster, then scored 20% lower on closed-book exams within six months. After about two years the exams that decide their school and university were down 24% and 18% (paper abstract). The more they used it, the worse it got.
Is an 18% rise in homework marks not a good thing?
It is a rise in the marks on work done with the tool in the room — not a rise in what the child can do without it. That is the whole finding: the same students went on to score 20% lower on closed-book exams, where the tool is not available (abstract).
Who ran it, and on whom?
David Strömberg of Stockholm University with Victor Lei and Yanhui Wu of the University of Hong Kong, following 26,811 students in grades 7 to 12, ages roughly 12 to 18, in one county in central China over 30 months (abstract). The SCMP dates the window September 2022 to June 2025 and reports that around 80% of them used generative AI at some point.
Has it been peer reviewed, and is one study enough?
No, and no single study ever is. It is a working paper on SSRN (abstract 6868618) and CEPR discussion paper DP21577 — released for other economists to attack, not yet accepted by a journal, and not a randomised trial. What makes it unusually strong anyway is the outcome measure: the real exams those children sat. Treat the mechanism as the finding and the exact percentages as local.
Was any group unharmed?
Close to it, and it is the most important detail. Students who used AI and still spent the same time on their homework came out with exam scores effectively identical to non-users' (PPC Land, on the paper's distribution figure); the abstract calls their losses "small". They also captured no time saving, which means the saving and the harm are largely the same thing.
Which children were hit hardest?
The strongest students, who lost around 24% against 16% for the bottom third, and younger children, who lost around 24% against 17% for older ones (The Decoder). A good report card is not evidence that nothing is happening.
What exactly did New York ban?
Student-facing generative AI from 2-K through eighth grade — about 600,000 children, two thirds of the system — for one year, starting with the 2026–27 school year, plus companion chatbots at every grade including high school. Teachers may still use AI for instructional planning and operational tasks, and there are exceptions for assistive technology, multilingual learners and career-readiness programmes (NYC Mayor's Office, 2 September 2026).
Did New York act because of this study?
Nothing says so. Neither the city's announcement nor Chalkbeat's reporting mentions it. The timing is close and the cuts line up with the evidence, but that is an observation, not a cause.
Does any of this bind a school in Germany?
No. New York City's moratorium is a local school-system policy with no effect here, and the study describes one Chinese county. If your school or employer is writing an AI policy, the transferable part is the mechanism, not the percentages — and the questions about pupil data, consent and works-council involvement belong to whoever owns that policy, not to a blog post.
Should I ban it at home too?
The study does not support abstinence — it supports order. Their attempt first, the tool afterwards. Being stuck after a real try and asking for help is how anyone learns; starting with the answer is the thing that did the damage.
How would I even know?
Ask your child to explain the work with the screen closed. Thin explanation, thin understanding. It takes ninety seconds and it is the only exam you have.
ISA After Hours · Augsburg
Working out what to do about this at home?
ISA After Hours is a community of Israeli and international tech professionals in Augsburg. Most of us have children in school here, we are all working this out at the same time, and we would rather do it together than each guess alone.
Join ISA After Hours →Sources
- scmp.comSouth China Morning Post: AI homework tools cut exam scores by 20%, study of 26,000 Chinese students finds
- papers.ssrn.comThe working paper itself — Strömberg, Lei and Wu, The Generative AI Learning Penalty: Evidence from Chinese Secondary Education, SSRN abstract 6868618 (also CEPR DP21577). SSRN blocks automated fetching, so the abstract was read via the two entries below.
- scale.stanford.eduStanford SCALE Initiative: the paper's abstract in full — 26,811 students, +18% homework, −20% monthly exams, −18 to −24% entrance exams
- ppc.landPPC Land: the paper's method and internal figures — the Callaway–Sant'Anna estimator, the standard deviations, and the distribution of time-matched users
- the-decoder.comThe Decoder: the breakdown by subject, age, usage level and prior attainment
- aljazeera.comAl Jazeera: Faster homework, poor exam results — what AI is doing to students' learning
- psychologytoday.comPsychology Today: A study of 26,000 students shows the AI learning trap
- kqed.orgKQED: "Metacognitive laziness" — how students offload critical thinking to AI
- nyc.govNew York City Mayor's Office: the generative AI moratorium in schools — scope, exceptions and the quotes from Mayor Mamdani and Chancellor Samuels
- chalkbeat.orgChalkbeat: NYC to ban student AI tools from 2-K through eighth grade and limit classroom screen time