The Winner's Curse in LLM Self-Improvement Loops: Selection Noise, Lock-in, and Acceptance Rules
We treat this keep-if-better step as selection under measurement noise, model the correlated errors of the candidates in a single decision, and study empirically what happens when the evaluation set is reused.
ProofPaper ↗
Key points
- Self-improving LLM systems propose changes to themselves and keep those that score better on a small evaluation set.
- In runs where Qwen models rewrite their own instructions and every candidate is also scored on 600 held-out items, most proposals after the first are harmful, and the model gives the size of the winner's curse of a generation's best candidate.
- With a prior from a separate pilot, it matches the average overstatement of first-generation commits in native loops, though not setting by setting.
- In a pre-registered study, the final selection-set score of greedy loops exceeded held-out accuracy by 13 to 20 points with 16 selection items and by 1 to 5 points with 256.
Sources (1)
- [1]The Winner's Curse in LLM Self-Improvement Loops: Selection Noise, Lock-in, and Acceptance RulesarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 12:05 AM
We treat this keep-if-better step as selection under measurement noise, model the correlated errors of the candidates in a single decision, and study empirically what happens when the evaluation set is reused.
Self-improving LLM systems propose changes to themselves and keep those that score better on a small evaluation set.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026Few Bits, One Law: Toward W2A4KV2
- Oct 6, 2026The Dichotomy Between Pattern Recognition and Step-by-Step Reasoning
- Oct 6, 2026Qwen3.8 Flash Next Reasoning Modes: Off vs Low vs Medium vs Xhigh
- Oct 1, 2026nvidia/PixelUMM
- Sep 5, 2026sgl-project/sglang v0.5.19
- Aug 22, 2026sgl-project/sglang v0.5.18