AION
Research paperReasoning & Planning · Applications · Safety & Alignment1 source · Oct 8, 2026

SWE-Journey: Towards More Realistic Evaluation of Coding Assistants through Long-Horizon, Multi-Turn Interaction

To address these gaps, we introduce SWE-Journey, a benchmark for more realistic evaluation of coding assistants.

Key points

  • Coding assistants such as Claude Code and Codex have become a major application of LLM agents, yet existing benchmarks remain far from real-world use, particularly in task horizon and interaction length.
  • Code assistants require completing long chains of development work in continuously evolving repositories, while repeatedly clarifying requirements and adapting implementations through multi-turn interaction.
  • To address the task-horizon gap, we propose a weak-to-strong synthesis pipeline that automatically constructs long-horizon coding tasks.
  • To address the interaction gap, we mine four representative user personas from real interaction data and build a user-simulation agent to reproduce realistic code-assistance interactions.

Sources (1)

Extractive summary: sentences quoted from the sources.