Budgeted Multi-Source Counterfactual Annotation for Off-Policy Evaluation
Off-policy evaluation (OPE) estimates the value of a target policy from logged data, but limited behavior-policy coverage can force high-variance reweighting or reward-model extrapolation.
ProofPaper ↗
Key points
- Counterfactual annotations can add evidence about unobserved actions, yet practical sources, including domain experts and large language models (LLMs), may be costly, biased, or noisy.
- We study budgeted acquisition of such annotations for contextual-bandit OPE.
- Given source-specific costs and error profiles, we formulate an integer allocation problem over context-action pairs and annotation sources to minimize the component of estimator variance that depends on the annotation plan.
- Experiments in synthetic clinical and LLM-annotated education bandits show that our allocation method reduces fixed-profile mean squared error (MSE) by 20.58% and 10.77%, respectively, relative to no annotation.
Sources (1)
- [1]Budgeted Multi-Source Counterfactual Annotation for Off-Policy EvaluationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 10:55 PM
Off-policy evaluation (OPE) estimates the value of a target policy from logged data, but limited behavior-policy coverage can force high-variance reweighting or reward-model extrapolation.
Counterfactual annotations can add evidence about unobserved actions, yet practical sources, including domain experts and large language models (LLMs), may be costly, biased, or noisy.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 7, 2026I would rather quit NLP than read another paper like this: The rise of antithesis in NLP papers
- Oct 7, 2026BoT-GRPO: Efficient Process-Reward RL for Reasoning via Bag-of-Token Aggregation
- Oct 7, 2026Judging in Latent Space: Efficient Generative Reward Modeling via Semantics-Preserving Compression
- Oct 6, 2026Personalize at Test Time: Learning User Preferences for Image Generation
- Oct 6, 2026Verify Less, Evolve More: Training Idea-Level Critics for Verification-Efficient ML Evolving Agents
- Oct 6, 2026Reinforcement Learning for Hierarchical Reasoning Rewards: Minimax-Optimal Rates with Transformers