Generative AI tools such as ChatGPT, Claude, and DeepSeek can help students solve problems, explain concepts, and polish their writing within seconds. This creates a productivity gain: students can complete assignments faster, with better scores. But in school, completing these tasks is not the goal; learning from them is. In a recent survey conducted by the Center for Democracy and Technology, more than 70% of parents and teachers reported concerns that students’ use of AI may weaken their academic skills (Laird and Dwyer 2025). Several OECD countries are developing guidance or policies to address the use of generative AI in education (OECD 2026). To guide policy, a clear understanding of the effects of generative AI on student learning is crucial. Are the effects positive or negative? How large are they? What students and subjects are more affected? What type of use is associated with what effects?
Recent studies regarding the effect of using specific AI tutors on student learning paint a mixed picture. A guarded, pedagogically designed AI tutor can improve learning (Kestin et al. 2025). Conversely, an AI tool that readily supplies answers improves students’ performance on practical questions (where AI can be used) but lowers scores on subsequent closed-book exams (Bastani et al. 2025). These studies provide useful evidence on how mandated use of AI tutors affects learning in narrowly defined curricular settings. However, much less is known about a broader question: how does students’ self-directed use of generative AI affect cumulative learning across subjects over time in natural education settings?
We study this question using 30 months of administrative records from 26,811 students in grades 7 to 12 in a county in central China (Strömberg et al. 2026). The data cover nine subjects and link weekly homework scores and completion time to monthly closed-book exams, county-wide exams, and high-school and college entrance examinations.
Importantly, we obtained students’ precise timing of generative AI adoption from surveys distributed by school teachers, who instructed students to check the registration dates for the AI tools they had used. By June 2025, roughly 80% of students reported using generative AI, up from almost none at the start of the sample in September 2022. The most commonly used AI tools were Doubao and DeepSeek, general-purpose AI tools without special educational functionality.
Because students began using AI at different times, we can compare how their outcomes changed after adoption with changes among students who had not adopted during the same period. Before adoption, the two groups had almost identical scores and followed similar trends. We use a staggered difference-in-differences design (Callaway and Sant’Anna 2021) to estimate the causal effects of students’ use of generative AI.
As shown in Figure 1, within six months of AI adoption, students, on average, completed a homework assignment in about 45 minutes, down from 64 minutes, which was the average before adoption. Their homework scores rose by 18% relative to the pre-adoption mean. AI indeed increased homework productivity: better output in less time.
But for learning, completing homework is not the end goal. Closed-book exam scores – used as a common measure of learning outcome – moved in the opposite direction. After six months, monthly exam scores fell by 20% relative to the pre-adoption average.
A penalty of a 20% fall in exam scores over six months is substantial. Similar large penalties are found in the randomised controlled trial by Bastani et al. (2025), who show that a generative AI tool providing direct answers to homework questions reduces closed-book math exam scores by 17%. Destructive learning effects of this magnitude are unusual among non-AI students in our sample.
Figure 1 Generative AI adoption changes homework scores, time spent on homework, and exam scores
For monthly exams, our estimated effect is small initially, then grows as students become more proficient with general-purpose tools and as more course material is learned after adoption. The full penalty appear about six months after students first adopted generative AI.
For the high-stakes entrance exams, the learning losses build more slowly over roughly two years. But when they reach their full magnitude, the penalty is roughly as large as that for regular exams: 24% of the baseline mean for high-school entrance examinations and 18% for the college entrance examinations. Since the entrance exams are the key determinant of school admissions, the losses have direct consequences for students’ educational trajectories and future careers.
The difference in how long it takes for the penalty to mature for the two types of exams is likely because regular exams mostly test recent material while entrance exams integrate material learned over several years. This implies that existing short-duration studies may systematically underestimate the long-run learning costs of using generative AI.
Traditionally, high homework scores indicate effective learning. Among non-AI students in our study, we see the expected positive relationship between homework scores and exam scores. Among AI users, however, this relationship changes dramatically. Those who score higher on homework are more likely to perform worse on exams. In this group, homework scores are no longer a useful metric to evaluate student learning. Instead, a rapid increase in homework scores sends a warning signal of learning decline. This suggests that, to restore learning, teachers and parents should monitor other aspects of learning, such as the time actually spent working on homework or performance on frequent, short, in-class quizzes to test homework material.
We observe two distinct ways of using generative AI, leading to divergent learning outcomes (Figure 2). After six months of use, around four out of every five AI-using students complete assignments in less than 50 minutes, faster than the fastest non-AI students, and receive exceptionally high homework scores (matching the capability of common AI tools). The exam scores of these efficient AI-using students, however, are markedly worse. This pattern is consistent with students outsourcing substantial parts of their homework to generative AI.
The remaining one out of five AI-using students continue to spend roughly as much time on homework as non-AI students do. These students receive similar exam scores, although they seem to use AI for homework, as evidenced by their higher homework scores. In other words, when AI students spend the same amount of time on homework as non-AI students, they learn as much.
This finding does not necessarily imply that learning will be restored if students are required to spend more time on homework. Students choosing different ways of using AI may also differ in motivation, information, and AI skills. Effective interventions need to combine supervised homework time with guidance on productive AI use and information about the learning losses due to outsourcing.
Figure 2 At equal homework time, AI users and non-users achieve similar exam scores
Our estimated AI learning penalty is broad but unevenly distributed. The largest estimated negative effect occurs in social-science subjects (e.g. history, politics), where exam scores fall by around 27%. The corresponding negative effects are about 22% in STEM subjects, 17% in English, and 9% in Chinese. Losses are disproportionately large for junior, male, and initially high-performing students.
Labour-market studies have found that AI compresses the skill distribution by disproportionately raising the productivity of less-skilled workers (e.g. Noy and Zhang 2023, Brynjolfsson et al. 2025, Cui et al. 2026) and university students (Hausman et al. 2025). In our educational setting, compression arises for the opposite reason: the learning penalty falls hardest on high performers.
Perhaps it is not surprising that students who outsource schoolwork to generative AI learn less. What is surprising, at least to us, is how large, pervasive, and persistent the learning losses are. Obviously, we cannot directly extrapolate effect sizes to other settings, but they should serve as a warning. The underlying features of the problem, such as the tendency of teenagers to avoid homework when they can, seems likely to be general.
Our findings suggest that the learning losses depend on how students use generative AI. In China and many other countries, well-designed tutoring tools already exist, often at nearly zero cost. Most students just do not use them. Rather, they prefer general-purpose AI tools that make it easy to outsource homework. Similar incentive problems are found in Oreopoulos and Low (2026), who show that students with access to an AI tutor configured to coach rather than give answers lack engagement with the prescribed AI tool.
This incentive problem will not go away with advancements in AI technology. The essential question in AI adoption in education is not technological, but organisational and institutional. More research should be devoted to address the issue of how to incentivise students to use AI tools in the right way in an environment where general-purpose AI tools provide direct answers.
Source: cepr.org
Falling agricultural prices have been found to generally increase armed civil conflict in producing regions,…
Geopolitical tensions have reshaped international trade and investment flows in recent years. This column uses…
Standard portfolio theory predicts strong investor responses to changes in the equity premium, but empirical…
Climate change and the transition to net zero are reshaping the macroeconomy, with important implications…
The role of non-bank financial intermediaries (NBFIs) in the financial system has increased markedly over…
Pension fund investors have been typically viewed as some of the most dependable investors in…