ResearchResearch paperReinforcement Learning1 source · Oct 6, 2026

Variance-Optimal Off-Policy Evaluation with Conjunct Effect Modeling

Off-policy evaluation (OPE) for contextual bandit policies becomes challenging when action-level importance weighting incurs excessive variance.

Key points

  • Doubly robust (DR) estimation remains unbiased under common support but retains these high-variance action-level weights.
  • A prior estimator, Off-policy evaluation with Conjunct Effect Model (OffCEM), replaces them with more stable cluster-level weights, at the cost of relying on local correctness of the reward model.
  • In this paper, we show that, under the assumptions required by DR and OffCEM, there exists an unbiased family of estimators that interpolates between OffCEM and DR.
  • Building on this result, we propose the Variance Optimal-CEM (VOCEM) estimator, which selects the interpolation coefficient to minimize variance.

Sources (1)

  • [1]Variance-Optimal Off-Policy Evaluation with Conjunct Effect Modeling
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 04:58 PM
    Off-policy evaluation (OPE) for contextual bandit policies becomes challenging when action-level importance weighting incurs excessive variance.
    Doubly robust (DR) estimation remains unbiased under common support but retains these high-variance action-level weights.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026Reinforcement Learning for Hierarchical Reasoning Rewards: Minimax-Optimal Rates with Transformers
  2. Oct 6, 2026SIGMA: Self-Improving Alignment Generalization from a Model Spec

Related