AION
Research paperReinforcement Learning1 source · Oct 7, 2026

CoTrace: Data Recipes for Training Terminal Agents with Harness-Model Co-Evolution

Within this framework, we introduce CoTrace, a harness-aware data recipe that explicitly governs trajectory routing, provenance matching, and curriculum refresh.

Key points

  • Terminal-agent capability depends jointly on model weights and the runtime harness that formats prompts, binds tools, and handles error recovery.
  • Existing harness-model co-evolution approaches improve both components, yet often treat trajectories produced during harness search as an undifferentiated replay buffer.
  • To systematically analyze this interface, we establish an alternating co-evolution framework that decouples harness search and policy training through component-wise promotion decisions.
  • Under CoTrace, recurring execution failures guide harness synthesis, while policy training is strictly conditioned on verified rollouts matched to the adopted runtime for supervised fine-tuning (SFT) or fresh online interactions for reinforcement learning (RL).

Sources (1)

  • [1]CoTrace: Data Recipes for Training Terminal Agents with Harness-Model Co-Evolution
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 05:06 PM
    Within this framework, we introduce CoTrace, a harness-aware data recipe that explicitly governs trajectory routing, provenance matching, and curriculum refresh.
    Terminal-agent capability depends jointly on model weights and the runtime harness that formats prompts, binds tools, and handles error recovery.

Extractive summary: sentences quoted from the sources.