Same tool. Opposite outcomes.
In the space of about a year, randomized studies found both that AI help during practice hurt students on the later unassisted exam, and that an AI tutor roughly doubledlearning gains. Same technology. Same subject areas. The difference wasn’t the model — it was whether the AI gave answers or hints.
16 min read · Includes an interactive: Mode Check
Not how much. How.
The intent hypothesis: what determines whether AI use erodes your capability is not exposure to AI but the mode you use it in.
Extractionis asking for the finished thing. The essay, the answer, the working code, the decision. You describe the problem and the output arrives complete. Your contribution is the specification and, if you’re careful, the check at the end.
Scaffolding is asking for help with an attempt you are already making. A hint toward the next step. A critique of your draft. An explanation of the error you just hit. The AI shapes the work; you still produce it.
These look almost identical from the outside. Same tool, same session length, often the same person within the same hour. The difference is invisible in the output and enormous in what it leaves behind in you.
problems solved unaided after ten minutes of AI-assisted work
Liu et al. 2026, preprint
problems skipped without any attempt, same study
Liu et al. 2026, preprint
learning gains from a tutor designed never to reveal solutions
Kestin et al. 2025
of student AI conversations are direct answer-seeking
Anthropic Education Report
What each study can actually claim.
These findings are not equally strong, and stacking them without saying so would be exactly the kind of thing this series exists to argue against.
Randomized · system-level · strongest
The high-school study: unrestricted AI help raised practice scores and lowered exam scores
Roughly a thousand students. Those given unrestricted GPT access during practice performed markedly better while they had it — and markedly worse than the control group on the later unassisted exam. Critically, a second version of the same tutor, restricted to hints rather than answers, eliminated the harm. The intervention was the system's design, assigned at random.
Randomized · system-level · strongest
The physics course: a never-reveal-solutions tutor roughly doubled learning gains
In a randomized comparison against active-learning classroom instruction, students working with an AI tutor built never to hand over the solution learned roughly twice as much, in less time. Same underlying technology as the study above. Opposite result. The design is the variable.
Randomized exposure · self-selected mode
The persistence trials: after ten minutes of AI help, people attempted less and solved less
Three studies, 1,222 participants. AI assistance was assigned at random; the AI was then removed. The assisted group solved fewer problems unaided and — the finding the authors flag as most important — skipped substantially more without attempting them at all. Within that group, people who had used the AI for hints showed no deficit while people who had extracted answers showed both effects. But they chose their own mode, so we cannot separate the mode from the person. Preprint, under review.
Preprint · small sample
Skill formation: developers learning with AI scored lower on later mastery
Developers learning an unfamiliar library with AI assistance scored substantially lower on a subsequent mastery test, with the largest gap in debugging — the part that most requires having built a mental model. Engaged interaction patterns preserved learning. Small sample, not yet peer-reviewed.
Usage data · descriptive
In the wild, roughly half of student AI use is direct answer-seeking
Not an experiment and not a claim about harm — just the base rate. If mode is the variable that matters, it is worth knowing which mode is the default.
How to read this stack
Taken together, these results are consistent with, not proof of, the intent hypothesis. The randomized evidence establishes that a system giving hints produces different outcomes than a system giving answers. It does not establish that a person choosing hints gets the same protection — that step is the leap, and it is currently supported only by data where people picked their own mode. Two of the studies above are preprints. One has a small sample. We are telling you this because a series about honest AI use that overstated its own evidence would be a bit rich.
The generation effect.
This part isn't new, isn't contested, and has nothing to do with AI. It's one of the most replicated findings in the study of learning.
Producing an answer yourself encodes it. Reading the same answer does not — or does so much more weakly. Psychologists have been demonstrating this since the 1970s with word pairs, and it has held up across essentially every kind of material anyone has tried it on. The effort of retrieving or constructing something is not an unfortunate cost of learning. It is the learning.
This is why the hint-versus-answer distinction does so much work. A hint preserves your generation attempt — you still have to produce the step. A complete answer converts the problem into a reading exercise. The task looks the same on your screen. Cognitively, one of them is the thing that builds capability and the other is a demonstration of it happening to someone else.
It also explains the shape of the persistence result. What dropped fastest wasn’t accuracy — it was attempting. As the researchers put it, AI “conditions people to expect immediate answers, thereby denying them the experience of working through challenges on their own.” Once you stop generating, there is nothing left for the effect to act on.
What nobody has shown yet.
No study randomizes intent
Every strong result here randomizes what the system does, not what the user is trying to do. Somebody needs to assign people to extraction and scaffolding conditions and see whether the effect survives. Until then the central claim of this guide is a hypothesis with good support, not a finding.
Cumulative real-world erosion is conjecture
The measured effects are immediate — same session, same day. Nobody has followed people for two years to see whether extraction habits compound into durable skill loss. It is plausible. It is not demonstrated.
The effects that exist are small to moderate
Thirteen percentage points on an unaided problem set is a real effect and not a catastrophe. Treat this as a reason to choose deliberately, not a reason to panic.
Self-selection runs through the user-level evidence
The people who ask for hints may simply be the people who were going to persist anyway. Disposition and mode are tangled together in every dataset we currently have on real users.
This is the hypothesis our field — and CPAI’s own research program — is testing. We would rather tell you that than sell you a certainty we don’t have.
The scaffolding moves.
Four phrasings that change the mode without changing the tool. Copy them, use them, adapt them.
Keeps you as the producer. You get the diagnosis; the repair is still your work, which is where the learning lives.
This is the phrasing that, built into a tutoring system, eliminated the harm in the randomized studies.
Errors are the highest-value learning moments you get. Handing over the fix spends that moment on nothing.
Also the antidote to sycophancy — models trained on human approval agree far more readily than an honest colleague would.
The stakes rule
Extraction is fine — genuinely fine — for low-stakes output you will never need to produce unaided. Reformat this table. Draft this scheduling email. Summarize this thread. Nobody is building a skill there and nobody needs to. Scaffolding is for anything you are trying to learn, or trying to keep. The mistake isn’t using extraction. The mistake is using it by default, on everything, without ever deciding which kind of task you’re looking at.
Which mode do you actually use?
Eight realistic moments. For each one, pick the prompt closest to what you’d genuinely type. Some of them should be extraction — the tool will say so.
Take the Mode Check →Action for every level of influence.
For yourself
- Before you open the chat, decide which this is: output you need, or a skill you're building. Extraction is the right call for the first. It is the wrong call for the second.
- Run one hints-only day. For anything you're trying to learn, ask for the next step rather than the finished thing, and notice how much longer it takes and how much more of it stays.
- When you catch yourself pasting in a problem and pasting out an answer, add one sentence: "Here's what I think it is — tell me where I'm wrong."
For knowledge workers
- Separate the deliverable from the capability. Nobody needs you to write meeting notes unaided. Somebody eventually needs you to reason through the thing the notes are about.
- In review, ask for the attempt as well as the output. A colleague who can show you their draft and the AI's critique of it has learned something the finished document doesn't record.
- Treat "I'd never be able to do this myself now" as data, not as a joke.
For educators
- Allow hints, prohibit answers — and say so as a rule about mode rather than a rule about tools. This is the intervention with the strongest randomized evidence behind it.
- Target both misconceptions at once. "Using AI to learn is cheating" and "any AI use is fine" are wrong for the same reason: both treat exposure as the variable.
- Adults need the workplace framing (deliverable vs. skill); students need the exam framing (the day the tool is gone).
For policy
- Require learning products to disclose hint-versus-answer design. In the randomized studies, that single design choice is the difference between harm and benefit.
- Fund the study nobody has run: randomize usage intent, not just system design, and follow users longer than a single session.
- Do not regulate AI in education by volume of use. The evidence does not support quantity as the variable.
Related
What You Hand Off
Offloading is old, useful, and has a price. The question isn't whether to hand things off — it's which capacities you're willing to let thin out.
The Perception Gap
Experienced developers using AI were 19% slower — while believing they were 20% faster. You cannot habit-correct what you cannot perceive.
Effort Is the Active Ingredient
Effortful, self-produced success is the ingredient both learning and mood regulation depend on. Frictionless help removes it.
Where this leads
Reading is one thing. Practicing it is another.
The Applied AI Certification builds practical AI fluency across all six domains — the working competence that advances toward proficiency, with structured practice, feedback, and a cohort on the same problems.
Research & further reading.
Want CPAI to teach this in your school or workplace?
We deliver the Healthy AI Use material as workshops and cohort programs for schools, libraries, employers, and community organizations.