SWE-bench
Also known as: SWE-bench Pro, SWE-bench Verified
14stories this week
14last 30 days
14all time
Timeline
- Oct 8, 2026 · Research paper · 1 sourceOnTrack: Real-Time Monitoring and Intervention in LLM Agent Trajectories via Streaming Structure-Aware Optimal TransportTo overcome this, we propose OnTrack, a streaming monitoring mechanism that compares an agent's steps and dependencies against recorded successful runs to alert users or block the agent in about a millisecond per step.
- Oct 8, 2026 · Research paper · 1 sourceAgentEvolver: System-Wide Self-Evolution Through Task ExecutionWe present AgentEvolver, a system for developing capabilities during task execution while keeping the foundation model fixed.
- Oct 8, 2026 · Research paper · 1 sourceChronos Enables Code Agents to Reason over Software EvolutionWe introduce Chronos, a test-time framework that makes this connected history available to large language model (LLM)-based code agents.
- Oct 8, 2026 · Research paper · 1 sourceOpera: A Verbal Critic Framework for Long-horizon Coding AgentsWe present Opera, a verbal critic framework that treats each correction as a persistent note, followed until the diagnosed problem is resolved.
- Oct 7, 2026 · Research paper · 1 sourceBefore They Can Solve: Predicting Post-Training Coding-Agent Performance from Base ModelsHow can we predict which base checkpoint is worth an expensive round of agentic post-training?
- Oct 7, 2026 · Research paper · 1 sourceCoTrace: Data Recipes for Training Terminal Agents with Harness-Model Co-EvolutionWithin this framework, we introduce CoTrace, a harness-aware data recipe that explicitly governs trajectory routing, provenance matching, and curriculum refresh.
- Oct 7, 2026 · Research paper · 1 sourceRSIGym: A Flexible Environment for Recursive Self-ImprovementWe introduce RSIGym, an agent-native research environment based on Everything as a Service (EaaS).
- Oct 7, 2026 · Product / feature launch · 1 sourceAgent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real HarnessesHarnessed Agentic RL: Microsoft Research Asia introduces a training paradigm in which the same agent harness used in deployment participates directly in reinforcement learning, removing the need to reimplement the agent inside the training framework.
- Oct 7, 2026 · Research paper · 1 sourceTestGRAD: Evolving Test Suites via Failure Pattern Momentum for SWE-Agent EnsembleSWE-agent ensembles improve issue resolution by combining candidate patches from different agents with complementary strengths.
- Oct 7, 2026 · Research paper · 1 sourceCoding-Agent Benchmarks Should Match Their Users' Task FlowsWe present SWE-TaskFlow, an approach for transforming any issue-derived benchmark: it preserves the verified tasks and tests while steering the interaction toward a target Task Flow through prompt splitting and verifiable repository QA, with a TaskFlow Alignment Score (TFAS) for selecting among generated trajectories.
- Oct 7, 2026 · Research paper · 1 sourceCode Understanding is a Bottleneck for Coding AgentsWe present CABRA: a Coding Ability Blueprint for Rigorous Agent evaluation.
- Oct 6, 2026 · Research paper · 1 sourceSpecGuard: Proving a Task Is Broken Before the Agent CheatsWe present SpecGuard, which detects and formally certifies these conflicts between task intent and tests.
- Oct 6, 2026 · Research paper · 1 sourceRippleCP: Measuring Counterfactual Checkpoint Advantage in Coding AgentsAgent checkpoint systems decide what state is recovery-relevant, how to snapshot it, and whether rollback is admissible.
- Oct 6, 2026 · Research paper · 1 sourceFC-SWE: Failure-Conditioned RL for Long-Horizon Software Engineering AgentsWe introduce FC-SWE, a failure-conditioned RL framework that incorporates recovery attempts into policy training.