AION
Research paperRobotics & Embodied AI1 source · Oct 8, 2026

AffordDrive3D: Affordance-Aware World-Action Modeling with Spatial Understanding

World-action models have recently improved autonomous driving by jointly learning future scene prediction and trajectory generation.

Key points

  • Jointly modeling action-relevant regions and future geometry can provide the policy with both driving-relevant cues and their corresponding spatial structure.
  • We therefore propose AffordDrive3D, an affordance- and geometry-aware world-action model that jointly learns future action-relevant regions and spatial structure.
  • In order to capture the scene semantics and driving context needed for driving affordance prediction, we build AffordDrive3D on a VLM backbone to forecast drivable areas and collision-critical regions that directly affect ego motion, while predicting future geometry from RGB world-model latents.
  • On NAVSIM, AffordDrive3D achieves state-of-the-art performance with 91.3 PDMS and 89.9 EPDMS, demonstrating the effectiveness of jointly modeling future affordances and geometry for trajectory planning.

Sources (1)

  • [1]AffordDrive3D: Affordance-Aware World-Action Modeling with Spatial Understanding
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 01:16 AM
    World-action models have recently improved autonomous driving by jointly learning future scene prediction and trajectory generation.
    Jointly modeling action-relevant regions and future geometry can provide the policy with both driving-relevant cues and their corresponding spatial structure.

Extractive summary: sentences quoted from the sources.