Constrained Command-Conditioned Reinforcement Learning with Bandit Strategy Selection in Real-Time Strategy Games
Deep reinforcement learning agents reach strong performance in real-time strategy games but can be brittle against opponents outside their training distribution.
Key points
- Separating strategic command selection from learned unit control allows different strategies to be selected for different opponents while reusing the same execution policy.
- We introduce a constrained command-conditioned Proximal Policy Optimization (PPO) policy, the executor, for MicroRTS, a real-time strategy environment.
- A Thompson-sampling bandit acts as the strategist, selecting command tuples from an estimate of the opponent's strategy built from in-game observations rather than opponent identity.
- In a controlled comparison with a flat PPO baseline trained with the same architecture, budget, curriculum and self-play league, the strategist-executor system wins significantly more often against three of the four strongest opponents on a training map, including the two strongest held-out ones (0.55 to 0.97 and 0.01 to 0.34), with no significant difference against the others.
Sources (1)
- [1]Constrained Command-Conditioned Reinforcement Learning with Bandit Strategy Selection in Real-Time Strategy GamesarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 10:37 AM
Deep reinforcement learning agents reach strong performance in real-time strategy games but can be brittle against opponents outside their training distribution.
Separating strategic command selection from learned unit control allows different strategies to be selected for different opponents while reusing the same execution policy.
Extractive summary: sentences quoted from the sources.