AION
Research paperReinforcement Learning1 source · Oct 7, 2026

BoT-GRPO: Efficient Process-Reward RL for Reasoning via Bag-of-Token Aggregation

We propose Bag-of-Tokens Group Relative Policy Optimization (BoT-GRPO), which extends GRPO to token-level reward models through a length-invariant "bag of tokens" aggregation: it collects all token-level rewards across rollouts, weights each by the inverse of its source sequence length, and computes per-token advantages relative to weighted group statistics.

Key points

  • Reinforcement learning is now central to eliciting reasoning in large language models, while in the popular algorithm Group Relative Policy Optimization (GRPO) every token in a rollout receives the same advantage.
  • We ask how to make process supervision efficient: accelerating convergence and improving final quality without the cost of value networks.
  • BoT-GRPO is critic-free, and is a drop-in replacement wherever GRPO is used when token-level reward is available.
  • On React front-end code generation, BoT-GRPO reaches $80%$ compile rate up to $1.9\times$ faster than GRPO and converges faster than modern GRPO variants (GSPO, DAPO, PURE) while reaching higher final compile and VLM-judged win rates.

Sources (1)

  • [1]BoT-GRPO: Efficient Process-Reward RL for Reasoning via Bag-of-Token Aggregation
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 10:23 AM
    We propose Bag-of-Tokens Group Relative Policy Optimization (BoT-GRPO), which extends GRPO to token-level reward models through a length-invariant "bag of tokens" aggregation: it collects all token-level rewards across rollouts, weights each by the inverse of its source sequence length, and computes per-token advantages relative to weighted group statistics.
    Reinforcement learning is now central to eliciting reasoning in large language models, while in the popular algorithm Group Relative Policy Optimization (GRPO) every token in a rollout receives the same advantage.

Extractive summary: sentences quoted from the sources.