World models
Also known as: world model
62stories this week
64last 30 days
66all time
Timeline
- Oct 8, 2026 · Research paper · 2 sourcesDreamTrue: Action-Faithful Robot World Model with Counterfactual Post-TrainingWe present DreamTrue, a multi-view, cross-embodiment robot world model for action-faithful and physically plausible video prediction.
- Oct 8, 2026 · Research paper · 1 sourceWhat 30,000 Hours of Ego-centric Video Does Not TeachWe show that the agent gains need not come from data, and a careful visual conditioning design saturates fidelity with a fraction of it, which lets us measure object fidelity on its own and discover its saturation point.
- Oct 8, 2026 · Research paper · 2 sourcesOuroWorld: Bringing Any 3D World Alive as Diverse, Endlessly Looping 3D CinemagraphsOuroWorld is a mask-free framework that turns any static 3D Gaussian Splatting scene into a 3D cinemagraph: a dynamic scene with vivid, diverse motion looping seamlessly from any viewpoint.
- Oct 8, 2026 · Research paper · 2 sourcesWorldGuide: Goal-Directed Video World Model for Procedural Task ExecutionWe formulate procedural video generation as closed-loop task execution in visual world space and introduce WorldGuide.
- Oct 8, 2026 · Research paper · 1 sourceWOVEN: Weaving Visual World Modeling into Multimodal LLMsWe therefore introduce WOVEN, a training source and benchmark for visual transition reasoning that organizes transition supervision by scene, action, and reasoning type, using diverse, realistic rollouts from video-pretrained generative models: 36,076 examples across 20 scene types, 5 action types, and 8 reasoning types.
- Oct 8, 2026 · Research paper · 1 sourceWorldCast: Distributed Multiplayer World ModelsWe present WorldCast, a distributed multiplayer world model in which each player runs a local client comprising a video generator and a state model.
- Oct 8, 2026 · Research paper · 1 sourceLeWAM: A JEPA World Action Model with Diffusion-Steering-Based MPCWe introduce LeWAM, a bidirectional transformer for forward, backward, inverse dynamics and policy prediction, on a decoder-free JEPA latent trained end-to-end through all four modes.
- Oct 8, 2026 · Research paper · 1 sourceLiteNWM: Efficient Latent World Models for Onboard Visual Navigation in the WildWe present LiteNWM, a latent navigation world model that shares visual encoding across candidates and jointly predicts their action-conditioned future representations at multiple horizons, while a learned scorer uses these predictions to select trajectories.
- Oct 8, 2026 · Research paper · 1 sourceRiCo: Neural Simulation of Rigid-Body Interactions via Local Contact ReasoningMotivated by this observation, we introduce Rigid-body Contact Reasoning (RiCo), which represents interactions between objects through sparse neighborhoods of contact surface points.
- Oct 8, 2026 · Research paper · 2 sourcesMulti-Agent Egocentric World Model with Fine-Grained Embodied InteractionWe propose Multi-agent Egocentric World Model (ME-World), which jointly denoises multiple ego streams in a shared token sequence, conditions each stream on all agents' target-view poses, and grounds generation with shared environment memory.
- Oct 8, 2026 · Research paper · 1 sourcePLaW-VLA: Predictive Latent World Modeling for Vision-Language-Action PoliciesLearning to predict how the world evolves can provide vision-language-action (VLA) policies with predictive context for long-horizon control, but its effectiveness depends on what future representation is modeled and how it conditions action generation.
- Oct 8, 2026 · Research paper · 1 sourceLanguage Models as AI Research World ModelsAI research agents automate the cycle of proposing, implementing, and evaluating experiments, opening a path toward recursive self-improvement.
- Oct 8, 2026 · Research paper · 1 sourceCausalDreamer: Learning Predictive World Models with Latent DisentanglementWe propose CausalDreamer, which keeps the tokenizer frozen and re-encodes its latent into a factored representation of four groups along two axes: controllability, where only the two controllable groups receive the action, and reward relevance, learned by predicting the reward from the two reward-relevant groups.
- Oct 8, 2026 · Research paper · 1 sourceRight Screen, Wrong Transition: World Models as Verifiers for GUI AgentsWe argue that a world model meant for verification should instead predict in the space in which observations are encoded, and present LGWM, a decoder-free, action-conditioned world model that predicts the representation of the next screen directly, trained without semantic annotation on 1.85M real GUI transitions.
- Oct 8, 2026 · Research paper · 2 sourcesMemento 3: Model-Based Recursive Self-Improvement through Reflective RulebooksWe introduce Memento 3, building on the Memento series to enable frozen LLM agents to continually learn explicit world models through external memory.
- Oct 8, 2026 · Research paper · 1 sourceMultiWorldBench: Do Independently Controlled Views Describe One Shared World?We introduce MultiWorldBench, a diagnostic Minecraft benchmark containing 495 case configurations across seven task suites and ten capabilities, including independent control, cross-view motion, shared-state synchronization, persistence, structural reasoning, concurrent interaction, and delayed revisit.
- Oct 8, 2026 · Research paper · 1 sourceActing from Belief, Looking When Needed: A Bayesian Spatial World Model for Navigation under Intermittent PerceptionWe study navigation under intermittent perception: acting from an internal spatial belief and looking again only when execution needs a new observation, potentially freeing the shared sensor for other tasks between navigation observations.
- Oct 8, 2026 · Research paper · 1 sourceSafe, Persistent, and Evolving Agent Harness for Understanding Partially Observable WorldsTo address these challenges, we introduce E-Ledger, a multi-agent harness for safe and persistent execution.
- Oct 8, 2026 · Research paper · 1 sourceAtomWorld-Mirror: Macro-Step World Modeling of Critical Evolution Backbones for Materials DynamicsWe propose AtomWorld-Mirror, a time-aware macro-step world model for the critical evolution backbone of atomic systems.
- Oct 8, 2026 · Research paper · 1 sourceLearning to Retrieve: Internalizing Memory Retrieval for Video World ModelsWe propose to internalize memory retrieval into the generation process, allowing retrieval to emerge as an intrinsic behavior of the video world model rather than relying on an external memory system.
- Oct 8, 2026 · Research paper · 1 sourceRewiring Semantics, Dynamics, and Control: A Simple yet Effective Action-Centric Tri-Stream TransformerVision-Language-Action (VLA) models have emerged as a prominent framework for complex robotic manipulation, building on the strong semantic understanding of pretrained Vision-Language Models (VLMs).
- Oct 8, 2026 · Research paper · 1 sourcePlanWAM: Planning-Shaped Future Representations for End-to-End Autonomous DrivingWorld models in end-to-end autonomous driving predict future scene evolution to provide foresight for trajectory planning.
- Oct 8, 2026 · Research paper · 1 sourceCRISP: Fixing Flying Pixels in Latent LiDAR Generation via Diffusion DecodingLatent LiDAR pipelines suffer from flying pixels: convolutional VAEs blur sharp radial depth discontinuities, yielding edge depths that back-project to points floating between surfaces.
- Oct 8, 2026 · Research paper · 1 sourceLLM-IDEA: Identifiability-Driven Experimental Agent for Autonomous Discovery of Mechanistic World ModelsWe propose the Identifiability-Driven Experimental Agent (LLM-IDEA) for closed-loop discovery with an identifiability engine that returns a three-way plateau verdict: capability limit, resolvable within the design class, or certified exhausted.
- Oct 8, 2026 · Research paper · 1 sourceVGGTWorld-VLA: Intent-Conditioned 3D World Evolution for Autonomous DrivingWe propose VGGTWorld-VLA, an intention-conditioned extension of VGGT-World for controllable 3D world evolution in autonomous driving.
- Oct 8, 2026 · Research paper · 1 sourceAffordDrive3D: Affordance-Aware World-Action Modeling with Spatial UnderstandingWorld-action models have recently improved autonomous driving by jointly learning future scene prediction and trajectory generation.
- Oct 7, 2026 · Research paper · 1 sourceWorld-Model Policy Arbiter for Goal-Conditioned Reinforcement LearningOffline goal-conditioned reinforcement learning (GCRL) has produced a diverse set of goal-reaching algorithms, yet no single algorithm performs best across environments, goals, and even different phases of the same task.
- Oct 7, 2026 · Research paper · 1 sourceCross-Embodiment Robot Foundation World Models with Latent ActionsWe introduce the Latent Action-Conditioned Robot World Model (LAC-WM), which operates within a learned unified latent action space shared across diverse embodiments.
- Oct 7, 2026 · Research paper · 1 sourceDiagnosing and Recovering from Observation-Space Shift at Long-Horizon Skill SeamsUsing privileged simulator resets, we find that the dominant shift comes from displaced scene state (e.g., an open drawer or secondary objects left behind by earlier skills), not from the robot's joint configuration or the object the downstream skill manipulates.
- Oct 7, 2026 · Research paper · 1 sourceMemoWM: How World Models Change What Agents Need to RememberWe formulate the problem of memory allocation conditioned on a world model and introduce MemoWM, a framework that uses shared predictions to compress retained information and reconstruct omitted content.