AION
Research paperReinforcement Learning1 source · Oct 8, 2026

Constrained Command-Conditioned Reinforcement Learning with Bandit Strategy Selection in Real-Time Strategy Games

Deep reinforcement learning agents reach strong performance in real-time strategy games but can be brittle against opponents outside their training distribution.

Key points

  • Separating strategic command selection from learned unit control allows different strategies to be selected for different opponents while reusing the same execution policy.
  • We introduce a constrained command-conditioned Proximal Policy Optimization (PPO) policy, the executor, for MicroRTS, a real-time strategy environment.
  • A Thompson-sampling bandit acts as the strategist, selecting command tuples from an estimate of the opponent's strategy built from in-game observations rather than opponent identity.
  • In a controlled comparison with a flat PPO baseline trained with the same architecture, budget, curriculum and self-play league, the strategist-executor system wins significantly more often against three of the four strongest opponents on a training map, including the two strongest held-out ones (0.55 to 0.97 and 0.01 to 0.34), with no significant difference against the others.

Sources (1)

Extractive summary: sentences quoted from the sources.