98% more pull requests, 91% longer reviews, flat delivery

Cumulus clouds stacked over a Colorado mountain range above a dry grassland

A new preprint on arXiv puts a name to something I think a lot of engineering managers have felt this year. The author calls it the productivity-reliability paradox. Telemetry from more than 10,000 developers shows pull requests (PRs) up 98% after AI coding tools arrived, review times up 91%, and the delivery metrics that matter to the business essentially flat. Meanwhile the controlled studies keep finding real gains, 20% to 56%, on well-scoped tasks.

Both things are true. The individual got faster at the part of the job that was already the fastest part. The organization's throughput is set somewhere else, at review, at integration, at the point where a change has to be understood by someone who did not write it, and that point just got twice as much traffic.

The paper proposes a specification governance model, which is a formal way of saying: decide what you want before the code is written, and review against that. I think that is directionally right. My simpler version is that review capacity is now the scarce resource on a team, and it should be planned like one. Who reviews, how long they have, what they are checking for, and what gets automated out of their way are the questions that decide whether the 98% turns into anything.

Has your team's review queue changed since the coding tools arrived, and did anything else change with it?

Photo source: https://photos.robertstowe.com/colorado