CPAI Issue Brief · Education & Schools · Learning & cognition
AI and Student Learning
AI can raise a student's performance while lowering their learning — and the gap between the two is invisible from inside the classroom.
What’s happening
Students use AI for schoolwork at scale, and the measurable question is no longer whether they use it. It is what happens to learning when they do. The research converges on an uncomfortable finding: performance during the task and capability after it can move in opposite directions, and the difference is decided by whether AI removes the effort or supports it.
What the evidence shows
The clearest result comes from a randomized trial of roughly 1,000 students across four practice sessions. Unrestricted GPT-4 access raised in-session grades 48%. On the subsequent unassisted exam, those same students scored 17% lower than classmates who never had access. A second version of the same model — instructed to give hints rather than solutions, and loaded with common student mistakes and the matching feedback — raised in-session grades 127% and produced essentially no effect on the unassisted exam. The harm was removed by pedagogical design, not by a better model.
Two 2026 field experiments sharpen this into a practical finding about structure. A two-year randomized trial across 18 middle schools, using an AI tutor explicitly configured to coach rather than answer, raised math achievement roughly 0.06 to 0.08 standard deviations over a school year — gains the researchers describe as resembling the same practice platform without AI at all. Their explanation is engagement: students rarely used it as a tutor. A companion randomized experiment with more than 6,000 students found AI's clearest benefit appeared specifically after mistakes, improving next-attempt correctness and shortening the path back to a correct answer. Delayed-test gains appeared only where AI was embedded in a mastery structure that forced engagement at the point of error.
The reason this is hard to manage locally is that no one in the room can feel it. A preregistered experiment found that any AI involvement impaired participants' memory a week later of which ideas had been their own. Preregistered studies with 2,691 participants found people systematically underestimate how much they rely on AI and overestimate what it saves them. Learning science supplies the mechanism: the practice conditions that feel most fluent reliably produce the least durable learning, and AI is very good at removing exactly the friction the learning was made of. Students themselves register the concern more than administrators do — 67% of surveyed youth agreed that more student AI use will harm critical thinking, against 22% of district leaders.
unrestricted AI raised practice performance and lowered the later unassisted exam
Bastani et al. 2025, PNAS — randomized trial
a coach-configured AI tutor over a school year — resembling the same platform without AI
Two-year RCT, 18 middle schools, 2026 — working paper
youth vs. district leaders who think student AI use will harm critical thinking
RAND, 2025–2026 surveys
Where it reaches constituents
Students at every level, and the teachers and families who cannot see the trade-off from outside it. The cost shows up only once the tool is taken away, in what the student can still do without it. That is also why individual course-correction is unreliable here: the signal people would need in order to adjust is precisely the one the research says they do not receive.
The current legal & regulatory landscape
This is largely unregulated territory. Federal Executive Order 14277 (April 2025) directs existing discretionary funds toward AI education without new appropriation. State activity concentrates on advisory guidance for districts and on student-facing curriculum requirements. Standards for how AI learning products are designed or evaluated are almost entirely absent.
No federal or state requirement currently exists that a learning product disclose whether it is built to give hints or answers — the single design variable the randomized evidence identifies as decisive. The evidence base itself is also thin: a Stanford review screened more than 800 papers and found only 20 high-quality causal studies, most of them short-term and measured on practiced material.
Considerations policymakers are weighing
- ·Whether AI learning products disclose the usage mode they are designed for — hints versus answers — since that is the variable the randomized evidence identifies as decisive.
- ·Which outcome measures count in evaluating a learning tool: performance during use, or performance once the tool is removed.
- ·Whether existing health-and-wellbeing instruction requirements are a natural home for AI usage content, or whether it belongs in computer science.
- ·Support for longitudinal measurement, which barely exists — nearly all current studies are single-semester and measured on practiced material.
Listed as live debates, not recommendations. CPAI does not take a position on how these should be resolved.
This brief condenses a full, sourced public guide. The complete evidence and citations:
Healthy AI Use (5-guide series) →More in Education & Schools
- Safe AI UseThe same AI tool produces opposite outcomes depending on how it is used — and usage guidance is largely absent while adoption is near-universal.
- Preparing Educators to Teach AI UseTeachers are being asked to supervise a technology whose effect on learning depends entirely on how students use it — and the guidance they receive is thinnest for exactly that job.
Key sources
A nonpartisan resource
The Center for Practical AI is a nonpartisan 501(c)(3) nonprofit. We provide education, research, and analysis, and we offer briefings and testimony on request. We do not endorse candidates or lobby for or against specific legislation. Everything here describes the evidence and the current landscape — the policy choices are yours.
Want CPAI to brief your office on this?
We provide nonpartisan briefings, research summaries, and testimony on request.
Request a briefing →