AION
Benchmark

SWE-bench

Also known as: SWE-bench Pro, SWE-bench Verified

14stories this week
14last 30 days
14all time

Timeline

  1. Oct 8, 2026 · Research paper · 1 source
    OnTrack: Real-Time Monitoring and Intervention in LLM Agent Trajectories via Streaming Structure-Aware Optimal Transport
    To overcome this, we propose OnTrack, a streaming monitoring mechanism that compares an agent's steps and dependencies against recorded successful runs to alert users or block the agent in about a millisecond per step.
  2. Oct 8, 2026 · Research paper · 1 source
    AgentEvolver: System-Wide Self-Evolution Through Task Execution
    We present AgentEvolver, a system for developing capabilities during task execution while keeping the foundation model fixed.
  3. Oct 8, 2026 · Research paper · 1 source
    Chronos Enables Code Agents to Reason over Software Evolution
    We introduce Chronos, a test-time framework that makes this connected history available to large language model (LLM)-based code agents.
  4. Oct 8, 2026 · Research paper · 1 source
    Opera: A Verbal Critic Framework for Long-horizon Coding Agents
    We present Opera, a verbal critic framework that treats each correction as a persistent note, followed until the diagnosed problem is resolved.
  5. Oct 7, 2026 · Research paper · 1 source
    Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models
    How can we predict which base checkpoint is worth an expensive round of agentic post-training?
  6. Oct 7, 2026 · Research paper · 1 source
    CoTrace: Data Recipes for Training Terminal Agents with Harness-Model Co-Evolution
    Within this framework, we introduce CoTrace, a harness-aware data recipe that explicitly governs trajectory routing, provenance matching, and curriculum refresh.
  7. Oct 7, 2026 · Research paper · 1 source
    RSIGym: A Flexible Environment for Recursive Self-Improvement
    We introduce RSIGym, an agent-native research environment based on Everything as a Service (EaaS).
  8. Oct 7, 2026 · Product / feature launch · 1 source
    Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses
    Harnessed Agentic RL: Microsoft Research Asia introduces a training paradigm in which the same agent harness used in deployment participates directly in reinforcement learning, removing the need to reimplement the agent inside the training framework.
  9. Oct 7, 2026 · Research paper · 1 source
    TestGRAD: Evolving Test Suites via Failure Pattern Momentum for SWE-Agent Ensemble
    SWE-agent ensembles improve issue resolution by combining candidate patches from different agents with complementary strengths.
  10. Oct 7, 2026 · Research paper · 1 source
    Coding-Agent Benchmarks Should Match Their Users' Task Flows
    We present SWE-TaskFlow, an approach for transforming any issue-derived benchmark: it preserves the verified tasks and tests while steering the interaction toward a target Task Flow through prompt splitting and verifiable repository QA, with a TaskFlow Alignment Score (TFAS) for selecting among generated trajectories.
  11. Oct 7, 2026 · Research paper · 1 source
    Code Understanding is a Bottleneck for Coding Agents
    We present CABRA: a Coding Ability Blueprint for Rigorous Agent evaluation.
  12. Oct 6, 2026 · Research paper · 1 source
    SpecGuard: Proving a Task Is Broken Before the Agent Cheats
    We present SpecGuard, which detects and formally certifies these conflicts between task intent and tests.
  13. Oct 6, 2026 · Research paper · 1 source
    RippleCP: Measuring Counterfactual Checkpoint Advantage in Coding Agents
    Agent checkpoint systems decide what state is recovery-relevant, how to snapshot it, and whether rollback is admissible.
  14. Oct 6, 2026 · Research paper · 1 source
    FC-SWE: Failure-Conditioned RL for Long-Horizon Software Engineering Agents
    We introduce FC-SWE, a failure-conditioned RL framework that incorporates recovery attempts into policy training.

Often appears with