ResearchResearch paperReinforcement Learning1 source · Oct 7, 2026

m-Set Adversarial Bandits with Winner Feedback

We show upper and lower bounds on the regret of $m$-set adversarial bandits for different utilities (winner reward or sum of rewards) and feedback models (winner index, winner reward, sum of rewards, and their combinations).

Key points

  • By comparing to standard bounds for combinatorial and MNL bandits, our results reveal how subtle changes in the setting can have a dramatic impact on the learning rates.
  • Our main technical contributions are the information-theoretic lower bounds on the regret.
  • Experiments on synthetic data confirm our theoretical analyses.

Sources (1)

  • [1]m-Set Adversarial Bandits with Winner Feedback
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 02:10 PM
    We show upper and lower bounds on the regret of $m$-set adversarial bandits for different utilities (winner reward or sum of rewards) and feedback models (winner index, winner reward, sum of rewards, and their combinations).
    By comparing to standard bounds for combinatorial and MNL bandits, our results reveal how subtle changes in the setting can have a dramatic impact on the learning rates.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026World Models Dream of Success: Diagnosing and Repairing Failure Insensitivity in Robot World Models
  2. Oct 6, 2026Learning Transition Kernels of Jump-Diffusion Processes with Conditional Diffusion Models
  3. Oct 6, 2026Algorithmic Scratchpads and Curriculum Staging for Arithmetic Reasoning in Tiny Transformers
  4. Oct 6, 2026Build a voice travel concierge with Amazon Bedrock AgentCore, Managed Knowledge Base and Nova Sonic
  5. Oct 4, 2026SheetSage2: Coherent Lead-Sheet Transcription with Synthetic Supervision
  6. Oct 2, 2026Inside-Out AI: Rebuilding Airbnb Behind the Scenes and Across the Guest Experience

Related