19% slower. It felt 20% faster.
In a 2025 randomized trial, experienced open-source developers worked on their own established projects — some tasks with AI, some without. They took 19% longer with the AI. They believed it had made them about 20% faster. A gap of nearly forty points, in experts, on their own code. If they couldn’t feel it, the working assumption should be that you can’t either.
13 min read · Includes an interactive: Calibration Check · Small study (16 developers) — labeled throughout
Faster was a feeling.
Told as a story, because the shape of it is the whole point: the people best placed to notice were the ones who didn't.
The researchers recruited sixteen experienced developers and had them work through real tasks on open-source projects they already knew well — not a toy benchmark, their actual repositories. Tasks were randomly assigned to be done with or without AI assistance. Everything was timed.
The developers expected the AI to speed them up. Afterwards, they reported that it had — by around 20%. The stopwatch disagreed: with AI, they were about 19% slower. The time went into prompting, waiting, reviewing, and fixing output that looked right and wasn’t. None of that registered as friction. What registered was the feeling of momentum.
This is a small study, and its authors say so plainly — sixteen people is sixteen people. Take the exact percentages as illustrative rather than a natural constant. But the directionis the finding, and it is the direction that matters: perceived and actual performance came apart, in the expert direction you’d least expect, on work these people understood better than anyone.
slower with AI, measured
METR developer RCT, 2025
faster, is how it felt
Same study, self-report
experienced developers — small study, labeled
METR 2025
their own established projects, not a benchmark
METR 2025
The same gap, everywhere.
One striking study is an anecdote. What makes this a habit worth building is that the underlying pattern shows up across unrelated literatures.
The efficiency-gain illusion
People underestimate how much they use AI and overestimate what it saves them. The subjective sense of time saved runs systematically ahead of the measured reality.
Fluency is not learning
Decades of research on learning versus performance: the study conditions that feel most smooth and productive produce the least durable learning. Feeling fluent is close to uncorrelated with having learned — and AI is a fluency machine.
Automation complacency doesn't spare experts
Reliance on an automated aid reallocates attention away from the task it's assisting. The aviation and process-control literature is emphatic that experience and skill do not immunize against this — if anything, trusted automation invites more of it.
Memory hides the evidence.
It's not just that you can't feel it in the moment. The record you'd use to catch it afterwards is also biased in the same direction.
A week after doing a task with AI, people often can’t tell which ideas were theirs and which were the AI’s — a preregistered, peer-reviewed experiment found that any AI involvement degraded people’s memory of who came up with what, most strongly in mixed human-AI work. So your personal history reads back as more your own than it was. The record you’d use to notice a slow decline is written in ink that flatters you.
The most sobering datapoint in this area comes from medicine. Experienced endoscopists’ unassisted detection rate fell about six percentage points (28.4% to 22.4%) after their clinics adopted AI-assisted procedures — professionals at the top of their field, who by every account felt entirely fine. That study is observational and genuinely contested; published correspondence argues the effect could be confounded by case mix and attention rather than true skill loss, and we cite it with that caveat attached rather than as a settled fact. But it points at exactly the mechanism this guide is about: the decline that doesn’t announce itself.
Evidence note
The AI memory-gap finding is preregistered and peer-reviewed, but it measures source memory — whose idea was it — not general skill. The endoscopist deskilling study is observational and contested; treat it as suggestive of a mechanism, not a measured law. The developer RCT is small. None of these individually would carry the claim. The reason to take the pattern seriously is that they point the same way from completely different starting points.
Why this guide sits in the middle of the series
You can’t habit-correct what you can’t perceive.
The first guide said the mode matters. The second said the hand-off has a price. Both of those failures are silent— the whole difficulty is that neither announces itself while it’s happening. Extraction feels productive. Offloading feels efficient. Skill slippage feels like nothing at all.
Which makes metacognitive awareness — the deliberate practice of checking your own performance against something other than your feelings about it — the habit that makes every other habit in this series usable. Without it, the advice is unfalsifiable. With it, you have a gauge.
Build a gauge.
You can't fix the perception gap by trying harder to perceive. You fix it by measuring — occasionally, deliberately, against reality.
Periodic unaided reps
On the keep-sharp tasks, do the thing without AI now and then. Not as a rule, as a measurement — the reading is the point, whatever it says.
Before-and-after estimates
Predict the time and the quality before you use AI. Compare after. This trains calibration and catches the efficiency illusion in the act.
The monthly no-AI hour
One hour a month on something that matters, tool closed. Not virtue, not detox — a gauge check on a skill you'd rather not discover has rusted at the worst possible moment.
An external spot-check
Ask a colleague to look at something you did unaided. An outside read reaches what self-perception structurally can't.
Watch your own calibration slip in three minutes.
Three quick tasks, done here, no AI. Estimate first, then find out. The gap between your guess and the number is the same gap this guide is about — just small enough to see in one sitting.
Take the Calibration Check →Action for every level of influence.
For yourself
- Do periodic unaided reps on your keep-sharp tasks. This is the only reliable signal you get — your feelings about your own performance are not it.
- Predict before you use AI: how long, how good, how many you'll get right. Compare afterwards. The gap is the whole point, and it shrinks with practice.
- Keep a monthly no-AI hour on something that matters — as measurement, not as virtue. You're checking the gauge, not proving anything.
For knowledge workers
- Ask a colleague to spot-check something you produced unaided. An external read catches what self-perception structurally can't.
- Distrust the feeling of speed specifically. In the developer study, the people who felt fastest were the ones being slowed down — expertise did not protect them.
- Treat 'I've still got it' as a hypothesis to test, not a fact you have access to.
For organizations
- Don't measure AI's impact by asking people how much it helped. Self-report and reality came apart by 39 points in the study that actually timed both.
- Build in occasional unaided benchmarks for skills the organization depends on. You cannot manage a drift you're not measuring.
- Beware productivity theatre: perceived speed-up is easy to report and, on this evidence, easy to get wrong.
For educators
- Teach calibration as a subject-agnostic skill: estimate, then measure, in whatever you're already teaching.
- Use the fluency-versus-learning dissociation directly — students feel most fluent in the conditions that teach them least, and knowing that changes how they study.
- Have students track prediction accuracy across a unit. The improving calibration is itself the learning.
Related
Extraction vs. Scaffolding
The same tool, opposite outcomes. What decides whether AI builds your capability or erodes it isn't how much you use it — it's how.
What You Hand Off
Offloading is old, useful, and has a price. The question isn't whether to hand things off — it's which capacities you're willing to let thin out.
Effort Is the Active Ingredient
Effortful, self-produced success is the ingredient both learning and mood regulation depend on. Frictionless help removes it.
Where this leads
Self-perception can't be the only gauge.
This whole guide is about why your own read on your capability is unreliable. A free CPAI assessment across the six proficiency domains gives you an external one — the thing self-perception structurally can't provide. And the Applied AI Certification builds the skills it measures.
Research & further reading.
Want CPAI to teach this in your school or workplace?
We deliver the Healthy AI Use material as workshops and cohort programs for schools, libraries, employers, and community organizations.