AION
Research paperReinforcement Learning · Large Language Models1 source · Oct 7, 2026

Which Rollout Taught It That? BehaviorTrace and the Limits of Training-Data Attribution in Online RL

When reinforcement learning teaches a language model a new behavior, can we find the training rollouts that taught it?

Key points

  • And when an attribution method says it can, how do we know the answer is real?
  • We study both questions on online RL fine-tuning with GRPO, using a planted behavior with a known cause.
  • We release BehaviorTrace, an open evaluation harness that combines full-gradient sketching, the planted-behavior setup, and controls for gradient magnitude, fluency, headroom, and variation across seeds and generation draws.
  • At saturated checkpoints, model fluency predicts the behavior label at least as well as every gradient method we compared it with.

Sources (1)

Extractive summary: sentences quoted from the sources.