SWE-Journey: Towards More Realistic Evaluation of Coding Assistants through Long-Horizon, Multi-Turn Interaction
To address these gaps, we introduce SWE-Journey, a benchmark for more realistic evaluation of coding assistants.
Key points
- Coding assistants such as Claude Code and Codex have become a major application of LLM agents, yet existing benchmarks remain far from real-world use, particularly in task horizon and interaction length.
- Code assistants require completing long chains of development work in continuously evolving repositories, while repeatedly clarifying requirements and adapting implementations through multi-turn interaction.
- To address the task-horizon gap, we propose a weak-to-strong synthesis pipeline that automatically constructs long-horizon coding tasks.
- To address the interaction gap, we mine four representative user personas from real interaction data and build a user-simulation agent to reproduce realistic code-assistance interactions.
Sources (1)
- [1]SWE-Journey: Towards More Realistic Evaluation of Coding Assistants through Long-Horizon, Multi-Turn InteractionarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 09:21 AM
To address these gaps, we introduce SWE-Journey, a benchmark for more realistic evaluation of coding assistants.
Coding assistants such as Claude Code and Codex have become a major application of LLM agents, yet existing benchmarks remain far from real-world use, particularly in task horizon and interaction length.
Extractive summary: sentences quoted from the sources.