Every evaluation of an AI coding tool I have seen was out of date before it finished. The tool changed under it, the model behind the tool changed under that, and the number it produced described a version nobody was running by the time the slides were shown. I think these tools are unusually hard to evaluate, and I want to name the reasons before saying what I would do anyway.
The target moves monthly
A traditional tool evaluation assumes the tool holds still long enough to be measured. These tools do not. The models are replaced every few months, the agents around them every few weeks, and the behavior of the same product in March and June can differ more than two competing products did in the same quarter. Any comparison is a snapshot, and the snapshot ages fast.
The task mix dominates
Every serious study this year found the same thing: the task matters more than the tool. Documentation changes, tests, and configuration go well. Feature work in a mature codebase goes worse. Which means an evaluation's result is mostly a function of which tasks were chosen, and a team that picks representative tasks badly will pick the wrong tool well.
The easy metrics measure the wrong thing
Time to complete and lines produced are easy to measure and nearly meaningless, because the tools are built to make those numbers go up. The things that matter, whether the code is maintainable, whether the author understands it, whether review got slower, whether anything shipped faster, are slow to measure and show up months later. An evaluation that runs for three weeks sees only the easy numbers.
Novelty wears off
Engineers trying a new tool are more careful, more engaged, and more enthusiastic than they will be in month four. Some of the measured gain in any trial is that attention, and it does not persist. The honest evaluation would run long enough for the novelty to fade, which is longer than anyone budgets.
Nobody is neutral
The people evaluating have usually already formed a view. Some want the tool, some fear it, and the vendor wants the sale. I have not seen a tool evaluation on this subject that was not, in part, a search for support of a conclusion already held.
What I would do anyway
Given all that, I would stop trying to pick the winner and start deciding what a tool would have to change for the team to care. Not output. Something downstream: review time, defect rate, time to onboard, how long a feature takes end to end. One or two of those, measured before, with a number the team agrees would mean the tool helped.
Then run the tool against that for longer than feels comfortable, on the team's real work rather than a chosen sample, and accept that the answer will be about this tool at this time on this codebase, not a verdict on the category.
And expect to do it again. The tools will keep changing. I think evaluation here is not a decision but a practice, repeated, with the same few downstream numbers watched over time. The team that has those numbers can tell when a new version actually helps. The team that ran one bake-off in the spring is guessing by autumn.
Photo source: https://photos.robertstowe.com/northern-arizona
