In 42.7% of AI-generated test suites, no test could tell apart the competing readings of the requirement. The model picked one and said nothing. A new arXiv paper on underspecified choices in AI code generation studied 600 suites from three models. Its fix: name each open question, build two throwaway versions that differ only there, and keep an input only if they disagree. A person picks the right result, and that input becomes a runnable record of the decision.
The paper is blunt about the usual remedy. An acceptance example written first resolves nothing when both readings pass it. I think plenty of teams with test-first habits around AI tools are there now: the tests pass, and none of them asked the question that mattered.
My view is that an ambiguous requirement is a product decision, and AI coding tools have moved it from whoever wrote the ticket to the model. Code review will not catch it, because the code is correct for the reading it chose.
After years of reading pull requests against tickets, I have come to believe the costly bugs are the ones that do exactly what someone thought the ticket meant.
When an agent writes code against one of your tickets, who decides what the ticket meant, and would anyone know a decision was made?
Photo source: https://photos.robertstowe.com/bermuda

