FC-SWE: Failure-Conditioned RL for Long-Horizon Software Engineering Agents
We introduce FC-SWE, a failure-conditioned RL framework that incorporates recovery attempts into policy training.
Key points
- Repository-level software engineering (SWE) is a challenging long-horizon setting: agents must reason over extended interactions, use tools, and adapt to stateful environments.
- Recent work trains SWE agents with reinforcement learning methods such as Group Relative Policy Optimization (GRPO), which independently sample multiple trajectories per issue, test the resulting patches, and compare terminal rewards within a fixed group.
- After a patch fails verification, FC-SWE restores the repository to its original task state and uses the failed patch and verifier feedback as context for a recovery trajectory.
- On all 500 SWE-bench Verified tasks under a verifier-assisted protocol, FC-SWE with Qwen3.5-4B and SWE-agent achieves 41.7% Resolved@1 and 52.8% Resolved@2, compared with 38.9% and 48.5% for GRPO.
Sources (1)
- [1]FC-SWE: Failure-Conditioned RL for Long-Horizon Software Engineering AgentsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 07:44 AM
We introduce FC-SWE, a failure-conditioned RL framework that incorporates recovery attempts into policy training.
Repository-level software engineering (SWE) is a challenging long-horizon setting: agents must reason over extended interactions, use tools, and adapt to stateful environments.
Extractive summary: sentences quoted from the sources.