Groundability, Not Scale Alone: When Weak Reviewers Can Audit Strong Coding Agents
Junyu Guo, Shangding Gu, Ming Jin, Javad Lavaei
Abstract
Coding agents can return plausible patches that omit required behavior. These failures are hard to review because long traces and confident summaries often hide what was missed. We ask when a nominally weaker reviewer can reliably decide whether a patch solves its issue. We study 411 execution-labeled traces from three agents and 101 controlled cases. On 154 GPT-5.4 traces, structured but unchecked evidence raises both defect catch and over-rejection. We then provide official execution evidence as an upper-bound diagnostic. After choosing and freezing one of two formats per reviewer, five of six reviewers improve both rates on 122 held-out traces; two classify every trace correctly. Reviewer size is not a consistent predictor of quality. Because official tests are unavailable in deployment, we also evaluate a frozen cascade with patch-caused static errors and generated tests that first fail on the unpatched repository. On 121 scored held-out GPT-5.4 traces and 59 Gemini traces, its coverage is 0.89 and 0.86, risk is 0.33 and 0.26, catch is 0.76 and 0.80, and over-rejection is 0.66 and 0.67. Most false rejections occur when unresolved cases reach the reviewer. Official execution evidence shows the potential of weak review when decisive checks are available. Producing equally reliable checks without official tests remains the main bottleneck.
Create a lesson
Related papers
Detecting Inconsistencies in Model Specifications with LLM-as-Verifier Reasoning
Zichen Xie, Mrigank Pawagi, Lize Shao et al.
CONTRA: Discovering and Qualifying Behavior-Changing Questions for Selective Clarification in LLM Code Generation
Zheng Fang, Yongmin Li, Yichang Zhang et al.
Code Detectors Have a Half-Life: Obsolescence and Metric Illusions in LLM-Generated Code Detection
Alberick Euraste Djire
Architectural Degradation: How to Measure and to Remediate
Noman Ahmad, Ruoyu Su, Matteo Esposito et al.
Refactoring React Component Hierarchies to Eliminate Prop Drilling
Vangelis Gkinis, Vassilis E. Zafeiris
A Design Theory for AI-Assisted Software Development Derived from Christopher Alexander's Theory of Form
Chien-Tsun Chen, Yu Chin Cheng