ResearchResearch paperReinforcement Learning1 source · Oct 6, 2026

Independent Multi-Agent Reinforcement Learning with Counterfactual Semantic-Social World Models

Fully decentralized multi-agent reinforcement learning (MARL), also referred to as independent learning, requires each agent to learn and act using only its local information and experience, without a centralized critic or inter-agent communication.

Key points

  • We argue that agents can learn more effectively by prospectively comparing the consequences of candidate actions rather than diagnosing failures only from realized returns.
  • We introduce CASTLE (Counterfactual Action-conditioned Semantic Tokens for Local Execution in Decentralized MARL), an offline-training, online-in-context guidance framework with two complementary world models.
  • A Local Dynamics World Model, offline pre-trained over agents' local trajectories, summarizes the agent's local trajectory dynamics and partial observability, while a Semantic-Social World Model predicts compact short-horizon task and social consequences for each candidate ego action.
  • The latter is trained from counterfactual simulator rollouts that expose plausible teammate and opponent responses to alternative actions taken from the same logged rollout state.

Sources (1)

  • [1]Independent Multi-Agent Reinforcement Learning with Counterfactual Semantic-Social World Models
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 03:53 AM
    Fully decentralized multi-agent reinforcement learning (MARL), also referred to as independent learning, requires each agent to learn and act using only its local information and experience, without a centralized critic or inter-agent communication.
    We argue that agents can learn more effectively by prospectively comparing the consequences of candidate actions rather than diagnosing failures only from realized returns.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026SPW-Nav Streams Language-Guided Panoramic Video in Real Time
  2. Oct 6, 2026A GPU-Parallel Framework for Heterogeneous Multi-Task Reinforcement Learning
  3. Oct 5, 2026From Traces to Agentic Worlds: Agentic Language World Models for Interactive Environment Simulation
  4. Sep 29, 2026Introducing Quine: An AI research system designed for the complexity of biology
  5. Sep 29, 2026LongTake: Learning to Sustain Dynamics in Long-Horizon Video Generation
  6. Jun 9, 2026Powering the future of robotics in Europe

Related