Which Rollout Taught It That? BehaviorTrace and the Limits of Training-Data Attribution in Online RL
When reinforcement learning teaches a language model a new behavior, can we find the training rollouts that taught it?
Key points
- And when an attribution method says it can, how do we know the answer is real?
- We study both questions on online RL fine-tuning with GRPO, using a planted behavior with a known cause.
- We release BehaviorTrace, an open evaluation harness that combines full-gradient sketching, the planted-behavior setup, and controls for gradient magnitude, fluency, headroom, and variation across seeds and generation draws.
- At saturated checkpoints, model fluency predicts the behavior label at least as well as every gradient method we compared it with.
Sources (1)
- [1]Which Rollout Taught It That? BehaviorTrace and the Limits of Training-Data Attribution in Online RLarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 05:01 PM
When reinforcement learning teaches a language model a new behavior, can we find the training rollouts that taught it?
And when an attribution method says it can, how do we know the answer is real?
Extractive summary: sentences quoted from the sources.