ResearchResearch paperLarge Language Models · Reinforcement Learning1 source · Oct 8, 2026

How to post-train on a surrogate: Envelope sampling mitigates reward hacking

In this work, we propose envelope sampling, a theoretically-grounded method for judge recalibration that seeks to minimize an upper bound on the regret of the post-trained model under the assumption that the human reward and re-calibrated reward lie in an $L^2$ ball around the judge.

Key points

  • Large language models (LLMs) are commonly post-trained against LLM judges and other cheap surrogates because the true reward, such as human preference, is too expensive to query at scale.
  • This practice often leads to reward hacking, where reinforcement learning against a miscalibrated surrogate leads to undesirable side effects.
  • In this work, we study a setting in which a small number $n$ of model outputs are annotated with ground-truth labels (e.g., from expert review) and used to recalibrate the LLM judge before optimizing against it.
  • We give practical algorithms to sample from the envelope by rejection or by fine-tuning against a modified reward, and experiments on clinical note generation and on a controlled sycophancy task show that recalibrating on envelope samples mitigates reward hacking where recalibrating on base-model samples does not.

Sources (1)

  • [1]How to post-train on a surrogate: Envelope sampling mitigates reward hacking
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:48 AM
    In this work, we propose envelope sampling, a theoretically-grounded method for judge recalibration that seeks to minimize an upper bound on the regret of the post-trained model under the assumption that the human reward and re-calibrated reward lie in an $L^2$ ball around the judge.
    Large language models (LLMs) are commonly post-trained against LLM judges and other cheap surrogates because the true reward, such as human preference, is too expensive to query at scale.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 8, 2026Beyond Imitation: A Framework and Benchmark for LLM-Assisted Peer Review
  2. Oct 8, 2026SuperNav: An Agentic Navigation System for Any Task in Any Scene
  3. Oct 8, 2026VibeEdit: Image Editing with Canvas Instructions
  4. Oct 8, 2026SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models
  5. Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
  6. Oct 7, 2026Q-Learning with Scalar Adjoint Matching

Related