Code review used to lean on a quiet signal that nobody talked about: code that looked right usually was. A person who could produce clean, idiomatic, well-named code usually understood the problem, so a reviewer could spend less attention on the clean changes and more on the messy ones.
An agent produces code that looks right every time. Clean, idiomatic, well-named, and confidently explained, whether or not it does the right thing. The signal is gone, and I think review has to change to make up for it. Here is what I look for.
Does it match the ask, and only the ask
The first thing I check is scope. Agents are eager. Asked to fix a bug, they will fix the bug, tidy the surrounding function, rename a variable, and add a helper nobody asked for. Each change is defensible. Together they turn a ten-line review into a two-hundred-line one, and the actual fix is hidden somewhere inside.
I think a reviewer should hold agent-written changes to a stricter scope than human ones, not a looser one. If the change does more than the task, the extra parts are unreviewed by anyone, including the author.
Do the tests test anything
Agents write tests readily, which is good, and the tests often pass, which proves less than it seems. I look at what the test would catch. A test that asserts the function returns what it currently returns is a snapshot, not a test. A test that mocks the thing being tested is decoration. The question is whether the test would fail if the behavior were wrong, and that takes reading the test, not counting it.
What could the agent not see
An agent sees the repository. It does not see the shape of the production data, the configuration that differs between environments, the downstream system that chokes on nulls, or the reason a strange-looking check was added three years ago. I think the reviewer's highest-value contribution is now exactly this: the things that are not in the code. If a change touches something whose constraints live outside the repository, the reviewer is the only one who knows.
What was weakened or duplicated
Two patterns I have learned to look for. The first is a check that got loosened to make a test pass: a type widened, a validation removed, a timeout extended. The second is a new helper that does what an existing helper already did, because the agent did not find the existing one. Both are easy to miss in a large diff and both compound.
Does the explanation match the diff
Agents write good pull request descriptions, and I think that is a hazard. A description is the author's claim about the change. With a person, the claim and the code were produced by the same understanding. With an agent, they were produced by the same model, and a model can describe what it meant to do rather than what it did. I read the diff first and the description second, and when they disagree, the diff is the truth.
The author still owns it
None of this moves responsibility to the reviewer. The person who opened the pull request owns every line in it, however it was produced, and I think a team should say so plainly. The agent is a tool the author used. The review is a second pair of eyes on the author's work, as it always was. What changed is only what those eyes should be looking for.
Photo source: https://photos.robertstowe.com/colorado
