AION
Concept

World models

Also known as: world model

62stories this week
64last 30 days
66all time

Timeline

  1. Oct 8, 2026 · Research paper · 2 sources
    DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-Training
    We present DreamTrue, a multi-view, cross-embodiment robot world model for action-faithful and physically plausible video prediction.
  2. Oct 8, 2026 · Research paper · 1 source
    What 30,000 Hours of Ego-centric Video Does Not Teach
    We show that the agent gains need not come from data, and a careful visual conditioning design saturates fidelity with a fraction of it, which lets us measure object fidelity on its own and discover its saturation point.
  3. Oct 8, 2026 · Research paper · 2 sources
    OuroWorld: Bringing Any 3D World Alive as Diverse, Endlessly Looping 3D Cinemagraphs
    OuroWorld is a mask-free framework that turns any static 3D Gaussian Splatting scene into a 3D cinemagraph: a dynamic scene with vivid, diverse motion looping seamlessly from any viewpoint.
  4. Oct 8, 2026 · Research paper · 2 sources
    WorldGuide: Goal-Directed Video World Model for Procedural Task Execution
    We formulate procedural video generation as closed-loop task execution in visual world space and introduce WorldGuide.
  5. Oct 8, 2026 · Research paper · 1 source
    WOVEN: Weaving Visual World Modeling into Multimodal LLMs
    We therefore introduce WOVEN, a training source and benchmark for visual transition reasoning that organizes transition supervision by scene, action, and reasoning type, using diverse, realistic rollouts from video-pretrained generative models: 36,076 examples across 20 scene types, 5 action types, and 8 reasoning types.
  6. Oct 8, 2026 · Research paper · 1 source
    WorldCast: Distributed Multiplayer World Models
    We present WorldCast, a distributed multiplayer world model in which each player runs a local client comprising a video generator and a state model.
  7. Oct 8, 2026 · Research paper · 1 source
    LeWAM: A JEPA World Action Model with Diffusion-Steering-Based MPC
    We introduce LeWAM, a bidirectional transformer for forward, backward, inverse dynamics and policy prediction, on a decoder-free JEPA latent trained end-to-end through all four modes.
  8. Oct 8, 2026 · Research paper · 1 source
    LiteNWM: Efficient Latent World Models for Onboard Visual Navigation in the Wild
    We present LiteNWM, a latent navigation world model that shares visual encoding across candidates and jointly predicts their action-conditioned future representations at multiple horizons, while a learned scorer uses these predictions to select trajectories.
  9. Oct 8, 2026 · Research paper · 1 source
    RiCo: Neural Simulation of Rigid-Body Interactions via Local Contact Reasoning
    Motivated by this observation, we introduce Rigid-body Contact Reasoning (RiCo), which represents interactions between objects through sparse neighborhoods of contact surface points.
  10. Oct 8, 2026 · Research paper · 2 sources
    Multi-Agent Egocentric World Model with Fine-Grained Embodied Interaction
    We propose Multi-agent Egocentric World Model (ME-World), which jointly denoises multiple ego streams in a shared token sequence, conditions each stream on all agents' target-view poses, and grounds generation with shared environment memory.
  11. Oct 8, 2026 · Research paper · 1 source
    PLaW-VLA: Predictive Latent World Modeling for Vision-Language-Action Policies
    Learning to predict how the world evolves can provide vision-language-action (VLA) policies with predictive context for long-horizon control, but its effectiveness depends on what future representation is modeled and how it conditions action generation.
  12. Oct 8, 2026 · Research paper · 1 source
    Language Models as AI Research World Models
    AI research agents automate the cycle of proposing, implementing, and evaluating experiments, opening a path toward recursive self-improvement.
  13. Oct 8, 2026 · Research paper · 1 source
    CausalDreamer: Learning Predictive World Models with Latent Disentanglement
    We propose CausalDreamer, which keeps the tokenizer frozen and re-encodes its latent into a factored representation of four groups along two axes: controllability, where only the two controllable groups receive the action, and reward relevance, learned by predicting the reward from the two reward-relevant groups.
  14. Oct 8, 2026 · Research paper · 1 source
    Right Screen, Wrong Transition: World Models as Verifiers for GUI Agents
    We argue that a world model meant for verification should instead predict in the space in which observations are encoded, and present LGWM, a decoder-free, action-conditioned world model that predicts the representation of the next screen directly, trained without semantic annotation on 1.85M real GUI transitions.
  15. Oct 8, 2026 · Research paper · 2 sources
    Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks
    We introduce Memento 3, building on the Memento series to enable frozen LLM agents to continually learn explicit world models through external memory.
  16. Oct 8, 2026 · Research paper · 1 source
    MultiWorldBench: Do Independently Controlled Views Describe One Shared World?
    We introduce MultiWorldBench, a diagnostic Minecraft benchmark containing 495 case configurations across seven task suites and ten capabilities, including independent control, cross-view motion, shared-state synchronization, persistence, structural reasoning, concurrent interaction, and delayed revisit.
  17. Oct 8, 2026 · Research paper · 1 source
    Acting from Belief, Looking When Needed: A Bayesian Spatial World Model for Navigation under Intermittent Perception
    We study navigation under intermittent perception: acting from an internal spatial belief and looking again only when execution needs a new observation, potentially freeing the shared sensor for other tasks between navigation observations.
  18. Oct 8, 2026 · Research paper · 1 source
    Safe, Persistent, and Evolving Agent Harness for Understanding Partially Observable Worlds
    To address these challenges, we introduce E-Ledger, a multi-agent harness for safe and persistent execution.
  19. Oct 8, 2026 · Research paper · 1 source
    AtomWorld-Mirror: Macro-Step World Modeling of Critical Evolution Backbones for Materials Dynamics
    We propose AtomWorld-Mirror, a time-aware macro-step world model for the critical evolution backbone of atomic systems.
  20. Oct 8, 2026 · Research paper · 1 source
    Learning to Retrieve: Internalizing Memory Retrieval for Video World Models
    We propose to internalize memory retrieval into the generation process, allowing retrieval to emerge as an intrinsic behavior of the video world model rather than relying on an external memory system.
  21. Oct 8, 2026 · Research paper · 1 source
    Rewiring Semantics, Dynamics, and Control: A Simple yet Effective Action-Centric Tri-Stream Transformer
    Vision-Language-Action (VLA) models have emerged as a prominent framework for complex robotic manipulation, building on the strong semantic understanding of pretrained Vision-Language Models (VLMs).
  22. Oct 8, 2026 · Research paper · 1 source
    PlanWAM: Planning-Shaped Future Representations for End-to-End Autonomous Driving
    World models in end-to-end autonomous driving predict future scene evolution to provide foresight for trajectory planning.
  23. Oct 8, 2026 · Research paper · 1 source
    CRISP: Fixing Flying Pixels in Latent LiDAR Generation via Diffusion Decoding
    Latent LiDAR pipelines suffer from flying pixels: convolutional VAEs blur sharp radial depth discontinuities, yielding edge depths that back-project to points floating between surfaces.
  24. Oct 8, 2026 · Research paper · 1 source
    LLM-IDEA: Identifiability-Driven Experimental Agent for Autonomous Discovery of Mechanistic World Models
    We propose the Identifiability-Driven Experimental Agent (LLM-IDEA) for closed-loop discovery with an identifiability engine that returns a three-way plateau verdict: capability limit, resolvable within the design class, or certified exhausted.
  25. Oct 8, 2026 · Research paper · 1 source
    VGGTWorld-VLA: Intent-Conditioned 3D World Evolution for Autonomous Driving
    We propose VGGTWorld-VLA, an intention-conditioned extension of VGGT-World for controllable 3D world evolution in autonomous driving.
  26. Oct 8, 2026 · Research paper · 1 source
    AffordDrive3D: Affordance-Aware World-Action Modeling with Spatial Understanding
    World-action models have recently improved autonomous driving by jointly learning future scene prediction and trajectory generation.
  27. Oct 7, 2026 · Research paper · 1 source
    World-Model Policy Arbiter for Goal-Conditioned Reinforcement Learning
    Offline goal-conditioned reinforcement learning (GCRL) has produced a diverse set of goal-reaching algorithms, yet no single algorithm performs best across environments, goals, and even different phases of the same task.
  28. Oct 7, 2026 · Research paper · 1 source
    Cross-Embodiment Robot Foundation World Models with Latent Actions
    We introduce the Latent Action-Conditioned Robot World Model (LAC-WM), which operates within a learned unified latent action space shared across diverse embodiments.
  29. Oct 7, 2026 · Research paper · 1 source
    Diagnosing and Recovering from Observation-Space Shift at Long-Horizon Skill Seams
    Using privileged simulator resets, we find that the dominant shift comes from displaced scene state (e.g., an open drawer or secondary objects left behind by earlier skills), not from the robot's joint configuration or the object the downstream skill manipulates.
  30. Oct 7, 2026 · Research paper · 1 source
    MemoWM: How World Models Change What Agents Need to Remember
    We formulate the problem of memory allocation conditioned on a world model and introduce MemoWM, a framework that uses shared predictions to compress retained information and reconstruct omitted content.

Often appears with