Variance-Optimal Off-Policy Evaluation with Conjunct Effect Modeling
Off-policy evaluation (OPE) for contextual bandit policies becomes challenging when action-level importance weighting incurs excessive variance.
ProofPaper ↗
Key points
- Doubly robust (DR) estimation remains unbiased under common support but retains these high-variance action-level weights.
- A prior estimator, Off-policy evaluation with Conjunct Effect Model (OffCEM), replaces them with more stable cluster-level weights, at the cost of relying on local correctness of the reward model.
- In this paper, we show that, under the assumptions required by DR and OffCEM, there exists an unbiased family of estimators that interpolates between OffCEM and DR.
- Building on this result, we propose the Variance Optimal-CEM (VOCEM) estimator, which selects the interpolation coefficient to minimize variance.
Sources (1)
- [1]Variance-Optimal Off-Policy Evaluation with Conjunct Effect ModelingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 04:58 PM
Off-policy evaluation (OPE) for contextual bandit policies becomes challenging when action-level importance weighting incurs excessive variance.
Doubly robust (DR) estimation remains unbiased under common support but retains these high-variance action-level weights.
Extractive summary: sentences quoted from the sources.