AION
Research paperRobotics & Embodied AI1 source · Oct 8, 2026

VGGTWorld-VLA: Intent-Conditioned 3D World Evolution for Autonomous Driving

We propose VGGTWorld-VLA, an intention-conditioned extension of VGGT-World for controllable 3D world evolution in autonomous driving.

Key points

  • VGGT provides a strong foundation for geometry-centric world models by recovering unified 3D scene geometry from visual observations.
  • First, we introduce an action--semantic conditioning mechanism that injects complementary driving semantics and ego-motion representations into the future-token stream, enabling different future geometry predictions for the same observed scene under alternative ego actions.
  • Second, we develop a geometry--language--action bridge that adapts historical geometry, VLA semantic features, and maneuver and trajectory representations for joint conditioning of future geometry prediction.
  • These results demonstrate the potential of semantic and action conditioning for controllable VGGT-based world prediction in autonomous driving.

Sources (1)

  • [1]VGGTWorld-VLA: Intent-Conditioned 3D World Evolution for Autonomous Driving
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 03:20 AM
    We propose VGGTWorld-VLA, an intention-conditioned extension of VGGT-World for controllable 3D world evolution in autonomous driving.
    VGGT provides a strong foundation for geometry-centric world models by recovering unified 3D scene geometry from visual observations.

Extractive summary: sentences quoted from the sources.