AION
Research paperLarge Language Models · Safety & Alignment · Reinforcement Learning1 source · Oct 7, 2026

RH-Detect: A Unified Benchmark for Reward Hacking Detection

We present RH-Detect, a benchmark that combines reward-hacking-relevant subsets from eleven public datasets, comprising 92,761 rows and six behavior categories, into a common schema.

Key points

  • Reward hacking, where a model exploits an evaluation signal without completing the intended task, threatens the reliability of deployed language model systems.
  • On 5,021 open-ended evaluation units, each comprising a task prompt and a free-form model continuation, including multi-turn tool-use trajectories, we evaluate six off-the-shelf language models from five families as reward hacking detectors without additional training.
  • We find that different input formats have different effects across models.
  • Our results show that a single pooled score can conceal variation across data sources, detector inputs, and training procedures.

Sources (1)

  • [1]RH-Detect: A Unified Benchmark for Reward Hacking Detection
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 09:52 PM
    We present RH-Detect, a benchmark that combines reward-hacking-relevant subsets from eleven public datasets, comprising 92,761 rows and six behavior categories, into a common schema.
    Reward hacking, where a model exploits an evaluation signal without completing the intended task, threatens the reliability of deployed language model systems.

Extractive summary: sentences quoted from the sources.