AION
Research paperReinforcement Learning · Efficiency & Inference1 source · Oct 7, 2026

Efficient Best-of-N policy evaluation for inference-time alignment

Best-of-N (BoN) is a common inference-time alignment method that selects the highest-scoring response among N samples from a reference model.

Key points

  • Evaluating BoN policies from logged data is challenging under sample-only access because standard off-policy estimators require density ratios that depend on unavailable response likelihoods.
  • In this paper, we propose a sample-only framework for evaluating and selecting BoN policies without access to these likelihoods.
  • We show that the order-statistic structure of BoN allows the required density ratios to be expressed through score-rank probabilities that are estimable from samples alone.
  • We establish valid asymptotic inference even under reward estimator misspecification and prove the efficiency of our BoN-DR estimator.

Sources (1)

  • [1]Efficient Best-of-N policy evaluation for inference-time alignment
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 12:26 AM
    Best-of-N (BoN) is a common inference-time alignment method that selects the highest-scoring response among N samples from a reference model.
    Evaluating BoN policies from logged data is challenging under sample-only access because standard off-policy estimators require density ratios that depend on unavailable response likelihoods.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026Few Bits, One Law: Toward W2A4KV2
  2. Oct 6, 2026The Dichotomy Between Pattern Recognition and Step-by-Step Reasoning
  3. Oct 6, 2026Verify Less, Evolve More: Training Idea-Level Critics for Verification-Efficient ML Evolving Agents
  4. Oct 6, 2026Reinforcement Learning for Hierarchical Reasoning Rewards: Minimax-Optimal Rates with Transformers
  5. Sep 5, 2026sgl-project/sglang v0.5.19
  6. Aug 22, 2026sgl-project/sglang v0.5.18

Related