ResearchResearch paperSafety & Alignment · Reasoning & Planning · Reinforcement Learning1 source · Oct 7, 2026

TestGRAD: Evolving Test Suites via Failure Pattern Momentum for SWE-Agent Ensemble

SWE-agent ensembles improve issue resolution by combining candidate patches from different agents with complementary strengths.

Key points

  • The central problem is therefore test-based selection: generate tests, execute candidate patches, and identify the best patch.
  • We formulate this process as test-space optimization: evolving an executable repository test suite until it distinguishes competing patches.
  • Inspired by gradient descent with momentum, we introduce TestGRAD, a framework for automatic test optimization.
  • On SWE-bench Verified, TestGRAD achieves 84.2% Pass@1 with a 4-agent ensemble, outperforming the strongest baseline (80.6%) by an absolute improvement of 3.6 percentage points, while compressing failure-history context by over $100\times$.

Sources (1)

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 7, 2026Coding-Agent Benchmarks Should Match Their Users' Task Flows
  2. Oct 7, 2026Code Understanding is a Bottleneck for Coding Agents
  3. Oct 6, 2026FC-SWE: Failure-Conditioned RL for Long-Horizon Software Engineering Agents
  4. Oct 6, 2026[AINews] Reflection Beam - 501B-A23B American Open Model
  5. Oct 2, 2026Academia is for Ambition — Alex Zhang, MIT
  6. Sep 30, 2026SCLATE: A Substrate for Continual-Learning Agent Training and Evaluation

Related