Variance-Averse $n$-Step Offline Reinforcement Learning for Sparse Long-Horizon Environments
We propose VAN-Flow (Variance-Averse $n$-step Flow), a framework that promotes reliable actions in generative offline RL.
ProofPaper ↗
Key points
- Generative actors are transforming offline reinforcement learning (RL) by enabling expressive policy classes that model complex action distributions.
- However, this expressiveness also exposes a key challenge in heterogeneous datasets: generative policies can reproduce unreliable action modes whose return distributions exhibit high variance, occasionally yielding high returns by chance but lacking consistency.
- Consequently, maximizing the expected $Q$-value alone is insufficient for identifying reliable actions.
- VAN-Flow combines (i) a categorical distributional critic, (ii) a variance-averse expectation operator that smoothly reweights atom probabilities to favor actions with both high returns and low dispersion, and (iii) a flow-matching generative actor guided via rejection sampling.
Sources (1)
- [1]Variance-Averse $n$-Step Offline Reinforcement Learning for Sparse Long-Horizon EnvironmentsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 07:47 AM
We propose VAN-Flow (Variance-Averse $n$-step Flow), a framework that promotes reliable actions in generative offline RL.
Generative actors are transforming offline reinforcement learning (RL) by enabling expressive policy classes that model complex action distributions.
Extractive summary: sentences quoted from the sources.
