AION
Research paperReasoning & Planning · Reinforcement Learning · Large Language Models1 source · Oct 8, 2026

Balancing Reference Guidance and Free Generation in Trajectory Rollouts for Reasoning RL

A verified reference solution provides a correct trajectory for training a reasoning model.

Key points

  • Alternatively, a prefix of the reference can guide the model in generating a trajectory of its own.
  • We study this question through prefix continuation, where the model continues from a reference prefix and keeps the resulting trajectory if it passes verification, falling back to the reference otherwise.
  • The resulting Adaptive Reference Guidance (ARG) constructs correct trajectories within a fixed generation budget, and we apply it to all-failure groups in Group Relative Policy Optimization (GRPO).
  • Experiments on Qwen3-4B and Qwen3-8B across five mathematical reasoning benchmarks show that ARG achieves the highest aggregate pass@12 among the evaluated methods with competitive average sampled accuracy.

Sources (1)

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 7, 2026SPLIT-RL: Staged Perception-Language Reasoning Training with Claim-Level Advantages
  2. Oct 7, 2026On the Clock: Towards Punctual and Productive Time-Budgeted AI Agents
  3. Oct 7, 2026RECAST: Learning to Compute the Right Context through Adaptive Evidence Routing
  4. Oct 7, 2026BoT-GRPO: Efficient Process-Reward RL for Reasoning via Bag-of-Token Aggregation
  5. Oct 7, 2026Visual Jev Rewards: Reference-Bound Verification for Multi-Subject Image Generation
  6. Oct 6, 2026FC-SWE: Failure-Conditioned RL for Long-Horizon Software Engineering Agents

Related