AION
Research paperReinforcement Learning1 source · Oct 7, 2026

Amortized Off-Policy Evaluation for LLMs

To address this, we propose PFN-OPE, a prior-data fitted network that amortizes OPE across a distribution of contextual-bandit tasks.

Key points

  • Accurate evaluation is central to selecting which LLM to deploy, yet testing a candidate on live traffic exposes real users to an unvetted model.
  • Teams therefore evaluate candidates offline, on data produced by already-deployed models.
  • This is off-policy evaluation (OPE), and it faces two distribution shifts: as a model is updated in post-training, its responses diverge from the logged ones (policy shift), and the reward definition under which it is judged changes with business requirements (reward shift).
  • On HelpSteer2 and UltraFeedback with Qwen, Llama, and Gemma policies, PFN-OPE achieves 2.0 to 9.3 times lower error than the best baselines across all tested configurations in the reward-shifted settings.

Sources (1)

  • [1]Amortized Off-Policy Evaluation for LLMs
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 07:51 PM
    To address this, we propose PFN-OPE, a prior-data fitted network that amortizes OPE across a distribution of contextual-bandit tasks.
    Accurate evaluation is central to selecting which LLM to deploy, yet testing a candidate on live traffic exposes real users to an unvetted model.

Extractive summary: sentences quoted from the sources.