ResearchResearch paperReinforcement Learning1 source · Oct 6, 2026

Variance-Averse $n$-Step Offline Reinforcement Learning for Sparse Long-Horizon Environments

We propose VAN-Flow (Variance-Averse $n$-step Flow), a framework that promotes reliable actions in generative offline RL.

Key points

  • Generative actors are transforming offline reinforcement learning (RL) by enabling expressive policy classes that model complex action distributions.
  • However, this expressiveness also exposes a key challenge in heterogeneous datasets: generative policies can reproduce unreliable action modes whose return distributions exhibit high variance, occasionally yielding high returns by chance but lacking consistency.
  • Consequently, maximizing the expected $Q$-value alone is insufficient for identifying reliable actions.
  • VAN-Flow combines (i) a categorical distributional critic, (ii) a variance-averse expectation operator that smoothly reweights atom probabilities to favor actions with both high returns and low dispersion, and (iii) a flow-matching generative actor guided via rejection sampling.

Sources (1)

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026StairVLA: Stage-Aware Hierarchical Action Generation for Vision-Language-Action Models
  2. Oct 6, 2026ESP: Energy-Score Policy for One-Step Multimodal Action Generation

Related