Agent pull requests merge at 65%, humans at 85%

The red sandstone mass of Courthouse Butte rising over juniper scrub near Sedona under a pale sky

A group at Polytechnique Montréal analyzed 220,612 closed pull requests (PRs) across 489 Python repositories with at least a hundred stars, 9,428 of them opened by coding agents: Codex, Copilot, Devin, Cursor, and Claude Code. Human PRs merged about 85% of the time across every quarter studied. Agent PRs merged between 64% and 68%, depending on the quarter, with no strong trend.

The averages hide the useful part. By agent, estimated merge rates ran from 43% for Devin to 84% for Claude Code, with Codex at 73%. By task, agents were sent mostly to documentation, dependency management, and testing, where configuration-style work merged above 80%, while function implementation and large language model (LLM) integration work merged below 60%.

I think the finding is consistent with every other study this year: agents do well on narrow, well-structured work and worse on anything that needs context the repository does not contain. What is new is the scale, and the scale suggests what a team should track in its own PR data.

Merge rate by author type, split by task category, is the first. Review comments per PR by author type is the second, because it is where the reviewer cost shows up. And time from open to merge by author type is the third. A team with those three numbers knows where its agents help and where they are making work for reviewers, which is more than any published study can tell it.

Does your team's PR data distinguish agent authors from human ones yet?

Photo source: https://photos.robertstowe.com/northern-arizona