Model Evaluation and Threat Research (METR), the group whose early 2025 study found that experienced developers were 19% slower with AI tools while believing they were faster, has published an update. A new round from late 2025 with 57 developers across 143 repositories and more than 800 tasks showed mixed results. But the part they lead with is a problem with the study itself: developers increasingly declined to take part because they were unwilling to work without AI, and between 30% and 50% of those who did take part avoided submitting tasks they expected AI would finish quickly. The study is now, in their words, systematically missing the developers most optimistic about AI, and they are changing the design.
I think the selection problem is the finding. A large share of developers will no longer work without the tools, which says something about how the tools have become part of the work.
My view is that this matters for anyone measuring AI tools inside their own team, because the same bias applies. The engineers most convinced the tools help are the least willing to be in the comparison group. In-house numbers inherit METR's problem without METR's care about it.
How does your team measure whether these tools help, and who chose which tasks went into the measurement?
Photo source: https://photos.robertstowe.com/bermuda

