AION
Research paperAgents & Tool Use · Evaluation & Benchmarks · Safety & Alignment1 source · Oct 7, 2026

Coding-Agent Benchmarks Should Match Their Users' Task Flows

We present SWE-TaskFlow, an approach for transforming any issue-derived benchmark: it preserves the verified tasks and tests while steering the interaction toward a target Task Flow through prompt splitting and verifiable repository QA, with a TaskFlow Alignment Score (TFAS) for selecting among generated trajectories.

Key points

  • The evaluation of coding agents generally strives to be as realistic as possible.
  • In our study, we collect 4,782 agent sessions of real software engineers in JetBrains IDEs, which we call Production Sessions.
  • Long-session samples from three public interaction corpora exhibit markedly different Task Flows (the distributions of session lengths, task types, and type-to-type transitions), so no single interaction distribution is universally realistic: benchmarks should name a target use case and calibrate to measurements from it.
  • In a pilot on 700 SWE-Bench Pro tasks, solving the task sequentially in several steps approximately doubles agent cost without a stable change in resolve rate: the interaction protocol itself is an important dimension of evaluation.

Sources (1)

  • [1]Coding-Agent Benchmarks Should Match Their Users' Task Flows
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 08:08 AM
    We present SWE-TaskFlow, an approach for transforming any issue-derived benchmark: it preserves the verified tasks and tests while steering the interaction toward a target Task Flow through prompt splitting and verifiable repository QA, with a TaskFlow Alignment Score (TFAS) for selecting among generated trajectories.
    The evaluation of coding agents generally strives to be as realistic as possible.

Extractive summary: sentences quoted from the sources.