Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving
Vision-language-action (VLA) models have emerged as a promising paradigm for autonomous driving.
Key points
- However, existing VLA models still suffer from a fundamental mismatch: driving actions require precise 3D geometric cues, while visual-language understanding and reasoning are largely conducted in a 2D semantic space.
- In this paper, we propose GeoCoTDrive, an explicit geometric chain-of-thought framework that grounds geometry in a planning-oriented manner.
- To supervise this process, we introduce planning-relevant grounding, a new region-level grounding task that focuses on local spatial cues directly affecting ego planning decisions, and construct the PlanningGrounding dataset to endow VLAs with planning-oriented grounding capability.
- Experiments across multiple end-to-end autonomous driving benchmarks show that GeoCoTDrive consistently improves safety-critical planning performance, demonstrating the effectiveness of the explicit geometric chain-of-thought process for VLA-based planning.
Sources (1)
- [1]Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous DrivingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 04:47 PM
Vision-language-action (VLA) models have emerged as a promising paradigm for autonomous driving.
However, existing VLA models still suffer from a fundamental mismatch: driving actions require precise 3D geometric cues, while visual-language understanding and reasoning are largely conducted in a 2D semantic space.
Extractive summary: sentences quoted from the sources.