Skip to content
Rivl
8 October 2026AI10 min

AI assisted software development speed, measured properly

The only randomised trials on this subject found experienced developers were slower with AI, then found them faster a year later, and the researchers trust neither number much. That is a more useful starting point than any vendor figure.

Every proposal we write has a version of this question behind it: does using AI tooling make the build faster, and by how much. The honest answer is that almost nobody has measured it properly, the people who did measure it got a result that surprised them, and when they repeated the experiment the result changed direction and they still would not bet on it.

That sounds like an evasion. It is actually the most useful thing available, because it tells you where the uncertainty sits. The claim we are careful never to make is a percentage.

The sections of this note in order: the trial that found AI made developers slower, how the number reversed while the researchers got less confident, why the honest answer is a range rather than a number, where the speed actually comes from, the cost that does not show up in a completion time, what we tell clients who ask for a percentage, and the honest limits.
How this note is organised.

The trial that found AI made developers slower

In July 2025 METR published a randomised controlled trial on experienced open source developers. Sixteen developers, 246 real issues on repositories they already knew well, each task randomly assigned to AI allowed or AI disallowed. The tooling was Cursor Pro with Claude 3.5 and 3.7 Sonnet, which is what was current at the time.

Tasks took 19 percent longer when AI was allowed. The detail that makes the study worth reading rather than just citing is the perception gap: the developers predicted beforehand that AI would make them 24 percent faster, and afterwards still believed it had made them about 20 percent faster. They were wrong in the same direction before and after, by roughly the same margin. The full write up is in METR's own report on the early 2025 study, and the authors are careful to say they are not claiming their developers or repositories represent most software work.

That caveat matters more than the headline. Mature repositories the developer already knows well is close to the worst case for AI assistance, because the thing AI is best at is the part an expert in that codebase does not need.

Then the number reversed, and the researchers got less confident

METR re ran the design with late 2025 tooling and published an update in February 2026. Writing speedups as negative numbers, their estimate for the original developers was minus 18 percent, meaning tasks got faster, against the previous year's plus 19 percent slowdown. A fresh group of developers came out at minus 4 percent.

If you stop there you get the headline that went round, which was that the finding had been reversed. The confidence intervals say something different. Minus 18 percent carries a range of minus 38 to plus 9, and minus 4 percent a range of minus 15 to plus 9. Both intervals cross zero, so neither result establishes a direction at all.

StudyPoint estimate95% interval
Early 2025 tools, original developers19% slower with AI+2% to +39%
Late 2025 tools, same developers18% faster with AI-38% to +9%
Late 2025 tools, new developers4% faster with AI-15% to +9%
The three METR estimates side by side: 19 percent slower with early 2025 tools on an interval of plus 2 to plus 39 percent, 18 percent faster with late 2025 tools on minus 38 to plus 9 percent, and 4 percent faster for new developers on minus 15 to plus 9 percent.
Every interval after the first one crosses zero.

METR is blunter about the weaknesses than its coverage was. It names developers declining to participate because they do not want to work without AI, which it says likely biases its speedup estimate downwards. It names a pay cut from 150 dollars an hour to 50 as a second selection effect. It says its measurements of time spent per task are unreliable for developers running several AI agents at once. And it notes the true speedup could be much higher among the developers and tasks selected out of the experiment. Its own summary of the evidence is that the data is only very weak evidence for the size of the increase. You can read the figures and the caveats in METR's February 2026 update.

Why the honest answer is a range and not a number

Three readings survive all of this. The effect is real but small compared to what is claimed. It depends heavily on the work, the codebase and the developer. And self reported speedup is not evidence, because the one trial that measured both found people confidently wrong about their own throughput.

The last one is the practical finding. If your estimate of how much faster your team has become comes from asking them, you have the same instrument that produced a 20 percent error in a controlled setting.

Where the speed actually comes from, in our experience

We will separate what the trials support from what they do not. The trials measured task completion time on existing repositories. They did not measure the parts of a project where we see the largest differences, and we are not going to pretend a 19 percent figure transfers to them.

  • Unfamiliar code. Reading a system nobody on the team wrote is where assistance pays most, and it is the opposite of METR's setup.
  • Scaffolding and the boring middle. Forms, migrations, adapters, test fixtures. Low judgement, high volume, and where the wall clock saving is visible.
  • First drafts of a thing nobody has specified yet. Producing something concrete to argue with beats producing a document about it.
  • Not the hard decisions. Choosing the data model, deciding what the product does not do, finding why the thing is slow. No measured speedup and plenty of plausible wrong answers.

That last line is the one that decides project outcomes. The failures we are called in to fix are almost never caused by typing speed, and the pattern is set out in why software projects fail.

The cost that does not show up in a completion time

Finishing a task faster and leaving the codebase in a worse state is not a speedup, it is a loan. Generated code tends to be locally reasonable and globally repetitive: four variations on one function where one belonged, each fine in isolation. Nothing in a task timer catches that, and it arrives as slowness three months later when every change has to be made in four places.

This is an old problem with a new delivery mechanism, and the vocabulary for it is in what technical debt actually means. The mitigation is unglamorous: review AI output against the shape of the system rather than against whether it works, and treat a pull request nobody can explain as unfinished.

What we tell clients when they ask for a percentage

We will not quote one, and anybody who does is quoting a number no randomised trial supports. What we will say is which parts of a scope we expect to move and which we do not, and we hold the estimate for the hard parts at the same place AI tooling or not.

The pattern is the same one that applies to every AI claim we evaluate: the demo is the easy case and the value is decided by the hard case. We have made that argument about output nobody checks in AI document processing and where it still needs a person, and about scores that look like measurements in AI lead scoring and when the number is just a guess.

It is not only a software question either. Whether a given AI tool clears its own cost is the same calculation a small business makes about any of them, and the version of it written for a non technical reader is Khaled Badr's piece on whether a chatbot earns its place for a small business.

The honest limits of everything above

Two randomised trials, both from one research group, both on open source work, with sixteen and then a slightly different small group of developers. That is a thin evidence base to run an industry on. A separate randomised trial of 96 Google engineers found roughly 21 percent faster completion on one complex enterprise task, with a wide interval and an explicit warning against generalising, which is consistent with the picture but does not thicken it much.

So the correct posture is neither the slowdown headline nor the speedup headline. It is that the effect is modest, conditional, and smaller than the confidence with which people assert it, including developers describing their own work. If you want a number for your own team, the only way to get one is to measure it, and asking will not do.

Describe it. We build it.

Seven or twelve days, pay on delivery, a year of maintenance included. Bring the problem, not a spec.

Book a meeting

Read next