A review of the vibe coding research, and what it agrees on

A round yellow door set into a grassy hillside, framed by red and pink zinnias in the sun

A group led by Dominik Michels published a state of the art review of the research on vibe coding, the practice of producing software by describing it to a model and accepting what comes back. It is the first thing I have read that puts the contradictory results in one place.

The contradictions are stark. Peer-reviewed experiments find 26% more tasks completed per week. An independent randomized trial finds a 19% slowdown. Team telemetry finds code review time up 441%. The authors argue the disagreement mostly resolves once you control for how productivity was measured, what scope the task had, and how long the study ran. Their central conjecture is that the gains are real on new code and shrink or reverse on mature codebases, which would explain why a lab study and a production team see different worlds.

The risk catalog is the other half: security failures in deployed applications, quality degradation in telemetry, skill atrophy, unsettled copyright exposure, and weak fault detection.

I think a review like this is what a team should write its internal policy from. Not a vendor's case study, but a document that has read the literature and says where it agrees. The policy that falls out: use the tools freely on new, well-scoped code, require more review on mature code than before, treat security review as unchanged, and keep some work for people to learn on.

Is your team's AI policy based on what the research says, or on the last demo someone saw?

Photo source: https://photos.robertstowe.com/new-zealand