Balancing Reference Guidance and Free Generation in Trajectory Rollouts for Reasoning RL
A verified reference solution provides a correct trajectory for training a reasoning model.
Key points
- Alternatively, a prefix of the reference can guide the model in generating a trajectory of its own.
- We study this question through prefix continuation, where the model continues from a reference prefix and keeps the resulting trajectory if it passes verification, falling back to the reference otherwise.
- The resulting Adaptive Reference Guidance (ARG) constructs correct trajectories within a fixed generation budget, and we apply it to all-failure groups in Group Relative Policy Optimization (GRPO).
- Experiments on Qwen3-4B and Qwen3-8B across five mathematical reasoning benchmarks show that ARG achieves the highest aggregate pass@12 among the evaluated methods with competitive average sampled accuracy.
Sources (1)
- [1]Balancing Reference Guidance and Free Generation in Trajectory Rollouts for Reasoning RLarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 02:54 AM
A verified reference solution provides a correct trajectory for training a reasoning model.
Alternatively, a prefix of the reference can guide the model in generating a trajectory of its own.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 7, 2026SPLIT-RL: Staged Perception-Language Reasoning Training with Claim-Level Advantages
- Oct 7, 2026On the Clock: Towards Punctual and Productive Time-Budgeted AI Agents
- Oct 7, 2026RECAST: Learning to Compute the Right Context through Adaptive Evidence Routing
- Oct 7, 2026BoT-GRPO: Efficient Process-Reward RL for Reasoning via Bag-of-Token Aggregation
- Oct 7, 2026Visual Jev Rewards: Reference-Bound Verification for Multi-Subject Image Generation
- Oct 6, 2026FC-SWE: Failure-Conditioned RL for Long-Horizon Software Engineering Agents