Agent pull requests: task type matters more than the agent

A rocky Dolomite peak above a green meadow and a wooden farmhouse under a blue sky crossed with contrails

A group at University College London looked at 7,156 pull requests (PRs) opened on GitHub by five coding agents: OpenAI Codex, GitHub Copilot, Devin, Cursor, and Claude Code. Instead of asking which agent is best, they split the PRs by what the task was, and the task turned out to be the dominant factor.

Documentation changes were accepted 82.1% of the time. New features were accepted 66.1% of the time, a 16 point gap. Within a single agent the spread was wider still: Codex ranged from 59.6% to 88.6% across the nine task categories. Claude Code was accepted on 92.3% of documentation PRs and 72.6% of feature PRs. Cursor did best on bug fixes at 80.4%.

I think this is the most useful kind of result for anyone evaluating these tools. A headline acceptance rate for an agent is an average over a mix of tasks that happens to be whatever that agent's users asked for. Change the mix and the number moves, with no change in the tool.

So my view on bake-offs is that they have to be stratified by task before the numbers mean anything. Pick the handful of task types the team actually does, run each tool on the same set of each, and compare within a type. An average across the lot tells you about the sample, not the tool.

When you compared coding tools, did you split the results by what the task was?

Photo source: https://photos.robertstowe.com/dolomites