AffordDrive3D: Affordance-Aware World-Action Modeling with Spatial Understanding
World-action models have recently improved autonomous driving by jointly learning future scene prediction and trajectory generation.
Key points
- Jointly modeling action-relevant regions and future geometry can provide the policy with both driving-relevant cues and their corresponding spatial structure.
- We therefore propose AffordDrive3D, an affordance- and geometry-aware world-action model that jointly learns future action-relevant regions and spatial structure.
- In order to capture the scene semantics and driving context needed for driving affordance prediction, we build AffordDrive3D on a VLM backbone to forecast drivable areas and collision-critical regions that directly affect ego motion, while predicting future geometry from RGB world-model latents.
- On NAVSIM, AffordDrive3D achieves state-of-the-art performance with 91.3 PDMS and 89.9 EPDMS, demonstrating the effectiveness of jointly modeling future affordances and geometry for trajectory planning.
Sources (1)
- [1]AffordDrive3D: Affordance-Aware World-Action Modeling with Spatial UnderstandingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 01:16 AM
World-action models have recently improved autonomous driving by jointly learning future scene prediction and trajectory generation.
Jointly modeling action-relevant regions and future geometry can provide the policy with both driving-relevant cues and their corresponding spatial structure.
Extractive summary: sentences quoted from the sources.