TestGRAD: Evolving Test Suites via Failure Pattern Momentum for SWE-Agent Ensemble
Pengfei He, Jiayuan Zhou, Shaowei Wang, Ruiqi Pan
Abstract
SWE-agent ensembles improve issue resolution by combining candidate patches from different agents with complementary strengths. The central problem is therefore test-based selection: generate tests, execute candidate patches, and identify the best patch. We formulate this process as test-space optimization: evolving an executable repository test suite until it distinguishes competing patches. Existing test-generation methods are limited optimizers. They usually lack an explicit loss for ensemble selection, optimize through incomplete directions that mostly create new tests or delete old ones, and perform one-off generation without feedback from repeated failures. Inspired by gradient descent with momentum, we introduce TestGRAD, a framework for automatic test optimization. TestGRAD centers on three concepts. Differential loss gives the optimizer an explicit execution-defined target: useful tests should separate candidate patches by behavior. Full CRUD gradients expand the update direction from merely creating or deleting tests to reading existing test infrastructure, creating new tests, updating stale assertions, and deleting only obsolete tests. Failure Pattern Momentum mines frequent failure sequences from memory, allowing the optimizer to avoid repeated non-discriminative directions while compressing the failure-history context. On SWE-bench Verified, TestGRAD achieves 84.2% Pass@1 with a 4-agent ensemble, outperforming the strongest baseline (80.6%) by an absolute improvement of 3.6 percentage points, while compressing failure-history context by over 100×.
Create a lesson
Related papers
TaoD2C-Bench: Benchmarking MLLMs for Industrial UI Code Generation Beyond Visual Fidelity
Chengwei Shi, Yunnong Chen, Tingting Zhou et al.
When Sub-Agents Work in Parallel: The Promises and Pitfalls of Dynamic Concurrency in Long-Horizon Coding Tasks
Han Li, HanHaoNing Li, Ziqian Jiang et al.
Using Small Language Models to Reverse-Engineer Machine Learning Pipelines Structures
Nicolas Lacroix, Frederic Precioso, Mireille Blay-Fornarino et al.
QuSema: Detecting Silent Bugs in Quantum Libraries via Quantum-knowledge-enhanced Agents
Yujin Song, Kaining Zhang, Qixin Zhang et al.
Why Software Engineering Is Indispensable in the Age of Coding Agents
Alfonso Fuggetta
AdaT2: Adaptive Test Transformations for Black-Box Boundary Testing of Conversational Agents
Liting Lin, Boxi Yu, Qinghua Xu et al.