Skip to main content
AI Transformation — all deep dives Data

AI productivity, measured

"×10", "90% of code written by AI"… The gap between the marketing and the measured reality is enormous. What the rigorous studies show — and the real conditions for a genuine gain.

Alexandre ZuberAlexandre Zuber — MyoAppPublished January 20, 202610 min read · Sources cited
Real AI productivity for developers — MyoApp analysis

Perceived vs measured

In 2025, the METR institute ran a rare experiment: rigorously measuring experienced developers on real open-source repositories, with and without AI assistance. Counter-intuitive result: with AI, they were roughly 19% slower — while they themselves estimated they had been faster. The gap between perceived and measured productivity is the real lesson of the study.

This does not mean AI is useless. It means the gain is neither automatic nor universal — and that it needs to be measured rather than assumed.

What METR actually measured

The protocol deserves to be known, because it explains the strength of the result. METR recruited experienced developers, maintainers of open-source repositories they knew intimately — not beginners on unfamiliar code. Each task (real issues from their own projects) was randomly assigned with or without AI assistance, using early-2025 tooling. Actual completion time was measured, and developers were also asked to estimate their gain.

Result: they estimated they were roughly 20% faster with AI… and were actually roughly 19% slower. Nearly forty percentage points between perception and measurement. It is this protocol — randomized, on real work — that makes the number hard to dismiss.

The caveats, honestly: the sample is small, the context is specific (experts on their own code, where AI has the least to add), and tools evolve quickly. The study does not say "AI slows everyone down" — it says that the gain is not automatic, even for the profiles you would expect to benefit most.

Where the gap between perception and reality comes from

Why do seasoned professionals misjudge their own speed so dramatically? Three well-identified mechanisms:

  • Invisible time: writing the prompt, waiting, reading the suggestion, correcting it — each loop feels short, but their sum does not register as "work." Typing your own code does.
  • Verification overhead: a plausible suggestion must be reviewed more carefully than your own code, precisely because you do not know where it might be wrong. That vigilance is tiring and slows you down — silently.
  • The pleasure of delegation: delegating produces a feeling of efficiency, independent of outcome. It feels good, so it feels fast. The stopwatch has no feelings.

The lesson is not "beware of AI" but "beware of self-reporting": any equipment decision based on "we feel more productive" rests on the least reliable metric that exists.

Where AI actually helps

  • Repetitive code: boilerplate, tests, migrations, documentation — high volume, low ambiguity.
  • Discovery: understanding an unfamiliar codebase, exploring an API, comparing approaches.
  • Languages and domains outside core expertise: AI helps more when a developer works outside their comfort zone than within it.

And where it wastes time

  • On complex, project-specific code, reviewing and correcting plausible-but-wrong suggestions costs more than writing it yourself.
  • Without project context (conventions, architecture, history), AI proposes generic code that creates technical debt.
  • The "it's moving forward" effect masks the real time spent iterating on prompts.
The lever is not the tool. It is the context you give it.

The conditions for a real gain

Teams that genuinely gain in productivity share three traits: they give AI the project context (conventions, architecture, decision history), they measure the impact instead of declaring it, and they maintain systematic human review. It is a question of method and technical leadership — not a tool subscription. That is, in fact, one of the first workstreams for a fractional CTO: transforming individual, scattered AI use into a tooled, measured team practice.

How to measure in your team

You do not need to replicate METR's protocol to measure the effect of AI in your team. Three simple metrics, tracked over a few weeks, are enough to move past self-reporting:

  • Cycle time by task type: duration from pickup to production, split between repetitive tasks and complex business logic — that is where the difference shows.
  • Rework rate: proportion of AI-assisted code revised in review or corrected after the fact. A speed gain cancelled by rework is not a gain.
  • Before/after comparison at constant scope: on the same category of tickets, not on different sprints — otherwise you are measuring project weather, not the tool.

This setup fits in a team dashboard and avoids the two symmetric pitfalls: declarative enthusiasm ("we're going twice as fast") and principled rejection. Between the two, there are your numbers.

What this means for your team

If the gain depends on context and method, then the organizational question matters more than the tooling question. Three practical consequences:

The role shifts toward review and scoping. The more AI produces code, the more value concentrates in those who can judge that code: architecture, edge cases, security. Technical judgment becomes the bottleneck — not typing speed. That is also why seniority matters more, not less, in the age of AI.

Project context becomes an asset. Written conventions, documented architecture, decision history: everything that helps a new developer also helps the AI — and multiplies its yield. Teams that invest in this "context engineering" widen the gap.

Training does not happen by itself. Knowing how to decompose a task for AI, detecting a plausible-but-wrong suggestion, deciding when to take back the keyboard: these are skills that need to be learned. Leaving them to develop "naturally" produces exactly the situation METR measured.

Frequently asked

Why do developers overestimate the gain from AI?

Because AI produces code quickly and continuously: the feeling of making progress is constant. But the time spent reviewing, correcting, and iterating on suggestions is systematically underestimated. That is precisely the gap the METR study highlights between perceived and measured productivity.

Will AI replace developers?

The available measurements do not point in that direction: they show a powerful tool whose yield depends heavily on the expertise of the person using it. The role shifts — more review, scoping, and architecture, less typing — but technical judgment remains the bottleneck.

What realistic gains can a team expect?

It depends on task type and method. The clearest gains appear on repetitive code, documentation, tests, and exploring unfamiliar codebases. On complex business logic, the gain can be zero or even negative without project context provided to the AI. That is why measuring on your real tasks matters more than trusting vendor numbers.

SourcesMETR, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity" (2025) — figures attributed to their authors.
Alexandre Zuber, fondateur de MyoApp
Alexandre ZuberFounder — MyoApp · La Rochelle, France

Over 25 years of software craftsmanship. These deep dives come straight from the workshop — the one that builds our SaaS, apps, and AI foundations — not from a marketing department. When a number is quoted, so is its source.

Meet the team
MyoApp · La Rochelle

Want a measured gain ?

We scope where AI genuinely accelerates your teams — and where it wastes time. No hype, just measurements.

Talk to a fractional CTO
We really listen. The rest follows.