Making AI Tutoring Productive: Evidence from a Mastery-Based Math Practice Experiment
Large language models (LLMs) can make tutoring more scalable, but only if students use them to reason through mistakes rather than avoid effort. We study this question in a randomized field experiment with more than 6, 000 middle-school students in Hamilton County Schools using NUMI, a research-based computer-assisted learning (CAL) platform. Students were randomized to AI versus CAL-only support, mastery versus non-mastery progression, and one of two math topics. One week later, they completed a delayed assessment covering practiced and unpracticed material. AI students progressed more slowly and attempted fewer questions, but answered more accurately conditional on reaching an attempt. The clearest mechanism appeared after mistakes: AI improved next-attempt correctness and reduced attempts needed to return to a correct answer, while increasing time spent on each question receiving structured support. Mastery increased three-correct-in-a-row attainment but did not by itself improve delayed learning. The most encouraging delayed-test evidence appears when AI is embedded in the mastery workflow, with marginally significant gains concentrated on practiced Exercise 1 material. The results suggest that LLM tutoring can add value over standard CAL, but its value depends on structure that turns mistakes into productive learning moments.
-
-
Copy CitationPhilip Oreopoulos, Michael Liut, Alp Sungu, and Nina Low, "Making AI Tutoring Productive: Evidence from a Mastery-Based Math Practice Experiment," NBER Working Paper 35621 (2026), https://doi.org/10.3386/w35621.Download Citation
-