Multi-Agent Egocentric World Model with Fine-Grained Embodied Interaction
We propose Multi-agent Egocentric World Model (ME-World), which jointly denoises multiple ego streams in a shared token sequence, conditions each stream on all agents' target-view poses, and grounds generation with shared environment memory.
Key points
- Egocentric world models predict first-person observations conditioned on an agent's actions, but most focus on a single agent.
- Real embodied settings often involve multiple agents that act and interact within a shared environment.
- Existing multi-agent world models rely on coarse actions like locomotion, camera control, or discrete commands, leaving fine-grained embodied interactions underexplored.
- We train and evaluate on real and synthetic multi-agent data and introduce shared-world consistency metrics for environment, update, and identity consistency.
Sources (2)
- [1]Multi-Agent Egocentric World Model with Fine-Grained Embodied InteractionHugging Face Daily Papers · Oct 8, 12:00 AM
We propose Multi-agent Egocentric World Model (ME-World), which jointly denoises multiple ego streams in a shared token sequence, conditions each stream on all agents' target-view poses, and grounds generation with shared environment memory.
Egocentric world models predict first-person observations conditioned on an agent's actions, but most focus on a single agent.
- [2]Multi-Agent Egocentric World Model with Fine-Grained Embodied InteractionarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 04:50 PM · same content
Extractive summary: sentences quoted from the sources.