AION
Research paperReinforcement Learning · Evaluation & Benchmarks1 source · Oct 6, 2026

Holdout Best-of-N: Unbiased Evaluation and Its Cost

We study evaluation from a fixed matrix of $K$ independent scores per candidate for a policy that selects using $J$ fresh scores.

Key points

  • Reusing the scores that select a Best-of-$N$ winner can overstate its expected reward.
  • A single estimator based only on this matrix is exactly unbiased for expected judge reward under every independent, stable collection of candidate-specific score laws if and only if $J<K$, for every pool size $M\ge N\ge2$.
  • For independent Gaussian scores with common variance and fixed $M\ge N\ge2$, the unbiased minimax risk in this regime is of order $σ^2/\sqrt K$, attained by Holdout; allowing bias improves the rate to $σ^2/K$.
  • At fixed selector depth, cyclic evaluation of bounded scores has $O(K^{-1})$ risk uniformly in pool size.

Sources (1)

  • [1]Holdout Best-of-N: Unbiased Evaluation and Its Cost
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 05:26 PM
    We study evaluation from a fixed matrix of $K$ independent scores per candidate for a policy that selects using $J$ fresh scores.
    Reusing the scores that select a Best-of-$N$ winner can overstate its expected reward.

Extractive summary: sentences quoted from the sources.