RH-Detect: A Unified Benchmark for Reward Hacking Detection
We present RH-Detect, a benchmark that combines reward-hacking-relevant subsets from eleven public datasets, comprising 92,761 rows and six behavior categories, into a common schema.
Key points
- Reward hacking, where a model exploits an evaluation signal without completing the intended task, threatens the reliability of deployed language model systems.
- On 5,021 open-ended evaluation units, each comprising a task prompt and a free-form model continuation, including multi-turn tool-use trajectories, we evaluate six off-the-shelf language models from five families as reward hacking detectors without additional training.
- We find that different input formats have different effects across models.
- Our results show that a single pooled score can conceal variation across data sources, detector inputs, and training procedures.
Sources (1)
- [1]RH-Detect: A Unified Benchmark for Reward Hacking DetectionarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 09:52 PM
We present RH-Detect, a benchmark that combines reward-hacking-relevant subsets from eleven public datasets, comprising 92,761 rows and six behavior categories, into a common schema.
Reward hacking, where a model exploits an evaluation signal without completing the intended task, threatens the reliability of deployed language model systems.
Extractive summary: sentences quoted from the sources.