Holdout Best-of-N: Unbiased Evaluation and Its Cost
We study evaluation from a fixed matrix of $K$ independent scores per candidate for a policy that selects using $J$ fresh scores.
Key points
- Reusing the scores that select a Best-of-$N$ winner can overstate its expected reward.
- A single estimator based only on this matrix is exactly unbiased for expected judge reward under every independent, stable collection of candidate-specific score laws if and only if $J<K$, for every pool size $M\ge N\ge2$.
- For independent Gaussian scores with common variance and fixed $M\ge N\ge2$, the unbiased minimax risk in this regime is of order $σ^2/\sqrt K$, attained by Holdout; allowing bias improves the rate to $σ^2/K$.
- At fixed selector depth, cyclic evaluation of bounded scores has $O(K^{-1})$ risk uniformly in pool size.
Sources (1)
- [1]Holdout Best-of-N: Unbiased Evaluation and Its CostarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 05:26 PM
We study evaluation from a fixed matrix of $K$ independent scores per candidate for a policy that selects using $J$ fresh scores.
Reusing the scores that select a Best-of-$N$ winner can overstate its expected reward.
Extractive summary: sentences quoted from the sources.