Center for Practical AI
Healthy AI Use · Guide 3The enabling habit

19% slower. It felt 20% faster.

In a 2025 randomized trial, experienced open-source developers worked on their own established projects — some tasks with AI, some without. They took 19% longer with the AI. They believed it had made them about 20% faster. A gap of nearly forty points, in experts, on their own code. If they couldn’t feel it, the working assumption should be that you can’t either.

13 min read · Includes an interactive: Calibration Check · Small study (16 developers) — labeled throughout

The flagship finding

Faster was a feeling.

Told as a story, because the shape of it is the whole point: the people best placed to notice were the ones who didn't.

The researchers recruited sixteen experienced developers and had them work through real tasks on open-source projects they already knew well — not a toy benchmark, their actual repositories. Tasks were randomly assigned to be done with or without AI assistance. Everything was timed.

The developers expected the AI to speed them up. Afterwards, they reported that it had — by around 20%. The stopwatch disagreed: with AI, they were about 19% slower. The time went into prompting, waiting, reviewing, and fixing output that looked right and wasn’t. None of that registered as friction. What registered was the feeling of momentum.

This is a small study, and its authors say so plainly — sixteen people is sixteen people. Take the exact percentages as illustrative rather than a natural constant. But the directionis the finding, and it is the direction that matters: perceived and actual performance came apart, in the expert direction you’d least expect, on work these people understood better than anyone.

19%

slower with AI, measured

METR developer RCT, 2025

~20%

faster, is how it felt

Same study, self-report

16

experienced developers — small study, labeled

METR 2025

own repos

their own established projects, not a benchmark

METR 2025

Not just developers

The same gap, everywhere.

One striking study is an anecdote. What makes this a habit worth building is that the underlying pattern shows up across unrelated literatures.

1

The efficiency-gain illusion

People underestimate how much they use AI and overestimate what it saves them. The subjective sense of time saved runs systematically ahead of the measured reality.

2

Fluency is not learning

Decades of research on learning versus performance: the study conditions that feel most smooth and productive produce the least durable learning. Feeling fluent is close to uncorrelated with having learned — and AI is a fluency machine.

3

Automation complacency doesn't spare experts

Reliance on an automated aid reallocates attention away from the task it's assisting. The aviation and process-control literature is emphatic that experience and skill do not immunize against this — if anything, trusted automation invites more of it.

Why it stays hidden

Memory hides the evidence.

It's not just that you can't feel it in the moment. The record you'd use to catch it afterwards is also biased in the same direction.

A week after doing a task with AI, people often can’t tell which ideas were theirs and which were the AI’s — a preregistered, peer-reviewed experiment found that any AI involvement degraded people’s memory of who came up with what, most strongly in mixed human-AI work. So your personal history reads back as more your own than it was. The record you’d use to notice a slow decline is written in ink that flatters you.

The most sobering datapoint in this area comes from medicine. Experienced endoscopists’ unassisted detection rate fell about six percentage points (28.4% to 22.4%) after their clinics adopted AI-assisted procedures — professionals at the top of their field, who by every account felt entirely fine. That study is observational and genuinely contested; published correspondence argues the effect could be confounded by case mix and attention rather than true skill loss, and we cite it with that caveat attached rather than as a settled fact. But it points at exactly the mechanism this guide is about: the decline that doesn’t announce itself.

Evidence note

The AI memory-gap finding is preregistered and peer-reviewed, but it measures source memory — whose idea was it — not general skill. The endoscopist deskilling study is observational and contested; treat it as suggestive of a mechanism, not a measured law. The developer RCT is small. None of these individually would carry the claim. The reason to take the pattern seriously is that they point the same way from completely different starting points.

Why this guide sits in the middle of the series

You can’t habit-correct what you can’t perceive.

The first guide said the mode matters. The second said the hand-off has a price. Both of those failures are silent— the whole difficulty is that neither announces itself while it’s happening. Extraction feels productive. Offloading feels efficient. Skill slippage feels like nothing at all.

Which makes metacognitive awareness — the deliberate practice of checking your own performance against something other than your feelings about it — the habit that makes every other habit in this series usable. Without it, the advice is unfalsifiable. With it, you have a gauge.

Practice

Build a gauge.

You can't fix the perception gap by trying harder to perceive. You fix it by measuring — occasionally, deliberately, against reality.

1

Periodic unaided reps

On the keep-sharp tasks, do the thing without AI now and then. Not as a rule, as a measurement — the reading is the point, whatever it says.

2

Before-and-after estimates

Predict the time and the quality before you use AI. Compare after. This trains calibration and catches the efficiency illusion in the act.

3

The monthly no-AI hour

One hour a month on something that matters, tool closed. Not virtue, not detox — a gauge check on a skill you'd rather not discover has rusted at the worst possible moment.

4

An external spot-check

Ask a colleague to look at something you did unaided. An outside read reaches what self-perception structurally can't.

Interactive

Watch your own calibration slip in three minutes.

Three quick tasks, done here, no AI. Estimate first, then find out. The gap between your guess and the number is the same gap this guide is about — just small enough to see in one sitting.

Take the Calibration Check →
What you can do

Action for every level of influence.

1

For yourself

  • Do periodic unaided reps on your keep-sharp tasks. This is the only reliable signal you get — your feelings about your own performance are not it.
  • Predict before you use AI: how long, how good, how many you'll get right. Compare afterwards. The gap is the whole point, and it shrinks with practice.
  • Keep a monthly no-AI hour on something that matters — as measurement, not as virtue. You're checking the gauge, not proving anything.
2

For knowledge workers

  • Ask a colleague to spot-check something you produced unaided. An external read catches what self-perception structurally can't.
  • Distrust the feeling of speed specifically. In the developer study, the people who felt fastest were the ones being slowed down — expertise did not protect them.
  • Treat 'I've still got it' as a hypothesis to test, not a fact you have access to.
3

For organizations

  • Don't measure AI's impact by asking people how much it helped. Self-report and reality came apart by 39 points in the study that actually timed both.
  • Build in occasional unaided benchmarks for skills the organization depends on. You cannot manage a drift you're not measuring.
  • Beware productivity theatre: perceived speed-up is easy to report and, on this evidence, easy to get wrong.
4

For educators

  • Teach calibration as a subject-agnostic skill: estimate, then measure, in whatever you're already teaching.
  • Use the fluency-versus-learning dissociation directly — students feel most fluent in the conditions that teach them least, and knowing that changes how they study.
  • Have students track prediction accuracy across a unit. The improving calibration is itself the learning.

Where this leads

Self-perception can't be the only gauge.

This whole guide is about why your own read on your capability is unreliable. A free CPAI assessment across the six proficiency domains gives you an external one — the thing self-perception structurally can't provide. And the Applied AI Certification builds the skills it measures.

Sources

Research & further reading.

Randomized controlled trialMETR (2025)Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer ProductivityRandomized controlled trial: 16 experienced developers, 246 tasks on their own mature open-source repositories (avg. 5 years' prior experience). They took 19% longer with AI assistance — despite forecasting a 24% speedup beforehand, and still estimating a 20% speedup afterward. Small sample; the authors are explicit about that.
Preprint · not yet peer-reviewedYu, Cheng, Jabbar, Sucholutsky, Collins, Jurafsky & Hawkins (2026)The Efficiency-Gain IllusionThree preregistered studies, N=2,691, on AI reliance for cognitively simple tasks. People frequently chose AI even when it saved no meaningful time or effort, with two systematic miscalibrations: believing they use AI less than they do, and overestimating what it saves. Prior AI use also begat further AI use within a session — a self-reinforcing loop. Preprint.
Review of prior researchSoderstrom & Bjork (2015)Learning Versus Performance: An Integrative ReviewPerspectives on Psychological Science. Conditions that make practice feel fluent and productive often produce the least durable learning — and vice versa. You feel most effective precisely when you are learning least.
Review of prior researchParasuraman & Manzey (2010)Complacency and Bias in Human Use of AutomationDecades of automation research across aviation, process control, and medicine. Complacency is strongest when automation is perceived as highly reliable, load is high, and the task is complex — all three describe ordinary AI use.
Preregistered experiment · peer-reviewedZindulka, Goller, Fernandes, Welsch & Buschek (2025)The AI Memory Gap: Misremembering What We Created With AIPreregistered experiment (n=184), published at CHI 2026. Participants generated ideas with and without an LLM; one week later, any AI involvement impaired source memory — knowing whether an idea was theirs or the AI's — most strongly in mixed human-AI workflows. Measures source attribution, not content retention.
Observational studyBudzyń, Romańczyk et al. (2025)Endoscopist deskilling after exposure to AI-assisted colonoscopyThe Lancet Gastroenterology & Hepatology. Retrospective observational study: 19 experienced endoscopists (2,000+ colonoscopies each) at four Polish centers. Their adenoma detection rate in non-AI colonoscopies fell 6.0 percentage points (28.4% to 22.4%) after routine AI assistance was introduced. Contested: published correspondence in the same journal (December 2025) raises the observational design, temporal and case-mix confounding, and whether this reflects true skill loss or attentional complacency — read it alongside the study.Published critique (Lancet correspondence, Dec 2025)
Survey / correlationalLee, Sarkar, Tankelevitch et al. (2025)The Impact of Generative AI on Critical ThinkingCHI 2025, Microsoft Research and Carnegie Mellon. N=319 knowledge workers across 936 real work tasks: higher confidence in the AI predicted less critical-thinking effort; higher confidence in one's own skill predicted more.
Last reviewed: July 2026We review this page quarterly. Statistics in this category change rapidly.The developer RCT is a small study (N=16) and is labeled as such wherever its numbers appear. The endoscopist deskilling study is observational and contested; published critiques are linked. The AI memory-gap study is preregistered and peer-reviewed but measures source memory, not general skill.

Want CPAI to teach this in your school or workplace?

We deliver the Healthy AI Use material as workshops and cohort programs for schools, libraries, employers, and community organizations.