Coding-Agent Benchmarks Should Match Their Users' Task Flows
We present SWE-TaskFlow, an approach for transforming any issue-derived benchmark: it preserves the verified tasks and tests while steering the interaction toward a target Task Flow through prompt splitting and verifiable repository QA, with a TaskFlow Alignment Score (TFAS) for selecting among generated trajectories.
Key points
- The evaluation of coding agents generally strives to be as realistic as possible.
- In our study, we collect 4,782 agent sessions of real software engineers in JetBrains IDEs, which we call Production Sessions.
- Long-session samples from three public interaction corpora exhibit markedly different Task Flows (the distributions of session lengths, task types, and type-to-type transitions), so no single interaction distribution is universally realistic: benchmarks should name a target use case and calibrate to measurements from it.
- In a pilot on 700 SWE-Bench Pro tasks, solving the task sequentially in several steps approximately doubles agent cost without a stable change in resolve rate: the interaction protocol itself is an important dimension of evaluation.
Sources (1)
- [1]Coding-Agent Benchmarks Should Match Their Users' Task FlowsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 08:08 AM
We present SWE-TaskFlow, an approach for transforming any issue-derived benchmark: it preserves the verified tasks and tests while steering the interaction toward a target Task Flow through prompt splitting and verifiable repository QA, with a TaskFlow Alignment Score (TFAS) for selecting among generated trajectories.
The evaluation of coding agents generally strives to be as realistic as possible.
Extractive summary: sentences quoted from the sources.