TestGRAD: Evolving Test Suites via Failure Pattern Momentum for SWE-Agent Ensemble
SWE-agent ensembles improve issue resolution by combining candidate patches from different agents with complementary strengths.
ProofPaper ↗
Key points
- The central problem is therefore test-based selection: generate tests, execute candidate patches, and identify the best patch.
- We formulate this process as test-space optimization: evolving an executable repository test suite until it distinguishes competing patches.
- Inspired by gradient descent with momentum, we introduce TestGRAD, a framework for automatic test optimization.
- On SWE-bench Verified, TestGRAD achieves 84.2% Pass@1 with a 4-agent ensemble, outperforming the strongest baseline (80.6%) by an absolute improvement of 3.6 percentage points, while compressing failure-history context by over $100\times$.
Sources (1)
- [1]TestGRAD: Evolving Test Suites via Failure Pattern Momentum for SWE-Agent EnsemblearXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 03:25 PM
SWE-agent ensembles improve issue resolution by combining candidate patches from different agents with complementary strengths.
The central problem is therefore test-based selection: generate tests, execute candidate patches, and identify the best patch.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 7, 2026Coding-Agent Benchmarks Should Match Their Users' Task Flows
- Oct 7, 2026Code Understanding is a Bottleneck for Coding Agents
- Oct 6, 2026FC-SWE: Failure-Conditioned RL for Long-Horizon Software Engineering Agents
- Oct 6, 2026[AINews] Reflection Beam - 501B-A23B American Open Model
- Oct 2, 2026Academia is for Ambition — Alex Zhang, MIT
- Sep 30, 2026SCLATE: A Substrate for Continual-Learning Agent Training and Evaluation