AION
Research paperReinforcement Learning1 source · Oct 7, 2026

Beyond Reward Suppression: Near-Optimal Offline Attacks on Warm-Start Bandits with Bounded Rewards

Adversarial attacks on bandits aim to mislead a learner toward a target arm while keeping the attack cost small.

Key points

  • Existing attacks typically achieve this by suppressing non-target arms.
  • We study this gap through bounded offline attacks on warm-start bandits, where an attacker can inject only valid action-reward pairs into the warm-start history before deployment.
  • We show that target promotion is not merely a heuristic: when the target arm lies near the lower reward boundary, any order-optimal-cost attack against UCB that makes it selected in nearly all online rounds must allocate a nonvanishing fraction of its cost to the target arm.
  • We further extend the attack to Thompson Sampling, $ε$-greedy, and a broader class of bandit algorithms.

Sources (1)

Extractive summary: sentences quoted from the sources.