AION
Research paperReinforcement Learning · Robotics & Embodied AI · Large Language Models1 source · Oct 7, 2026

Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models

How can we predict which base checkpoint is worth an expensive round of agentic post-training?

Key points

  • End-to-end pass@$K$ tests whether successful behavior already appears in a base model's distribution, but it is a poor fit for agentic coding: many base checkpoints cannot reliably produce the well-formed tool invocation required to complete a task end-to-end.
  • Single-shot or short-horizon tasks avoid these tool-calling failures by collapsing a multi-step interaction into a fixed prompt and a single patch, but they sidestep the core capability we care about: maintaining coherent state over many tool-using steps as the repository evolves.
  • To bridge this gap, we treat successful post-trained agent trajectories as a lookahead signal of base-model potential.
  • As our methods need only a benchmark's successful trajectories and its verifier, they can be applied to turn future agentic coding benchmarks into base-model evaluations.

Sources (1)

  • [1]Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 05:37 PM
    How can we predict which base checkpoint is worth an expensive round of agentic post-training?
    End-to-end pass@$K$ tests whether successful behavior already appears in a base model's distribution, but it is a poor fit for agentic coding: many base checkpoints cannot reliably produce the well-formed tool invocation required to complete a task end-to-end.

Extractive summary: sentences quoted from the sources.