Documentation changes were accepted 82.1% of the time. New features were accepted 66.1% of the time, a 16 point gap. Within a single agent the spread was wider still: Codex ranged from 59.6% to 88.6% across the nine task categories. Claude Code was accepted on 92.3% of documentation PRs and 72.6% of feature PRs. Cursor did best on bug fixes at 80.4%.
I think this is the most useful kind of result for anyone evaluating these tools. A headline acceptance rate for an agent is an average over a mix of tasks that happens to be whatever that agent's users asked for. Change the mix and the number moves, with no change in the tool.
So my view on bake-offs is that they have to be stratified by task before the numbers mean anything. Pick the handful of task types the team actually does, run each tool on the same set of each, and compare within a type. An average across the lot tells you about the sample, not the tool.
When you compared coding tools, did you split the results by what the task was?