ResearchResearch paperReinforcement Learning · Large Language Models1 source · Oct 7, 2026

Budgeted Multi-Source Counterfactual Annotation for Off-Policy Evaluation

Off-policy evaluation (OPE) estimates the value of a target policy from logged data, but limited behavior-policy coverage can force high-variance reweighting or reward-model extrapolation.

Key points

  • Counterfactual annotations can add evidence about unobserved actions, yet practical sources, including domain experts and large language models (LLMs), may be costly, biased, or noisy.
  • We study budgeted acquisition of such annotations for contextual-bandit OPE.
  • Given source-specific costs and error profiles, we formulate an integer allocation problem over context-action pairs and annotation sources to minimize the component of estimator variance that depends on the annotation plan.
  • Experiments in synthetic clinical and LLM-annotated education bandits show that our allocation method reduces fixed-profile mean squared error (MSE) by 20.58% and 10.77%, respectively, relative to no annotation.

Sources (1)

  • [1]Budgeted Multi-Source Counterfactual Annotation for Off-Policy Evaluation
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 10:55 PM
    Off-policy evaluation (OPE) estimates the value of a target policy from logged data, but limited behavior-policy coverage can force high-variance reweighting or reward-model extrapolation.
    Counterfactual annotations can add evidence about unobserved actions, yet practical sources, including domain experts and large language models (LLMs), may be costly, biased, or noisy.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 7, 2026I would rather quit NLP than read another paper like this: The rise of antithesis in NLP papers
  2. Oct 7, 2026BoT-GRPO: Efficient Process-Reward RL for Reasoning via Bag-of-Token Aggregation
  3. Oct 7, 2026Judging in Latent Space: Efficient Generative Reward Modeling via Semantics-Preserving Compression
  4. Oct 6, 2026Personalize at Test Time: Learning User Preferences for Image Generation
  5. Oct 6, 2026Verify Less, Evolve More: Training Idea-Level Critics for Verification-Efficient ML Evolving Agents
  6. Oct 6, 2026Reinforcement Learning for Hierarchical Reasoning Rewards: Minimax-Optimal Rates with Transformers

Related