How to post-train on a surrogate: Envelope sampling mitigates reward hacking
In this work, we propose envelope sampling, a theoretically-grounded method for judge recalibration that seeks to minimize an upper bound on the regret of the post-trained model under the assumption that the human reward and re-calibrated reward lie in an $L^2$ ball around the judge.
ProofPaper ↗
Key points
- Large language models (LLMs) are commonly post-trained against LLM judges and other cheap surrogates because the true reward, such as human preference, is too expensive to query at scale.
- This practice often leads to reward hacking, where reinforcement learning against a miscalibrated surrogate leads to undesirable side effects.
- In this work, we study a setting in which a small number $n$ of model outputs are annotated with ground-truth labels (e.g., from expert review) and used to recalibrate the LLM judge before optimizing against it.
- We give practical algorithms to sample from the envelope by rejection or by fine-tuning against a modified reward, and experiments on clinical note generation and on a controlled sycophancy task show that recalibrating on envelope samples mitigates reward hacking where recalibrating on base-model samples does not.
Sources (1)
- [1]How to post-train on a surrogate: Envelope sampling mitigates reward hackingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:48 AM
In this work, we propose envelope sampling, a theoretically-grounded method for judge recalibration that seeks to minimize an upper bound on the regret of the post-trained model under the assumption that the human reward and re-calibrated reward lie in an $L^2$ ball around the judge.
Large language models (LLMs) are commonly post-trained against LLM judges and other cheap surrogates because the true reward, such as human preference, is too expensive to query at scale.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 8, 2026Beyond Imitation: A Framework and Benchmark for LLM-Assisted Peer Review
- Oct 8, 2026SuperNav: An Agentic Navigation System for Any Task in Any Scene
- Oct 8, 2026VibeEdit: Image Editing with Canvas Instructions
- Oct 8, 2026SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models
- Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
- Oct 7, 2026Q-Learning with Scalar Adjoint Matching