AION
Research paperRobotics & Embodied AI1 source · Oct 7, 2026

Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving

Vision-language-action (VLA) models have emerged as a promising paradigm for autonomous driving.

Key points

  • However, existing VLA models still suffer from a fundamental mismatch: driving actions require precise 3D geometric cues, while visual-language understanding and reasoning are largely conducted in a 2D semantic space.
  • In this paper, we propose GeoCoTDrive, an explicit geometric chain-of-thought framework that grounds geometry in a planning-oriented manner.
  • To supervise this process, we introduce planning-relevant grounding, a new region-level grounding task that focuses on local spatial cues directly affecting ego planning decisions, and construct the PlanningGrounding dataset to endow VLAs with planning-oriented grounding capability.
  • Experiments across multiple end-to-end autonomous driving benchmarks show that GeoCoTDrive consistently improves safety-critical planning performance, demonstrating the effectiveness of the explicit geometric chain-of-thought process for VLA-based planning.

Sources (1)

  • [1]Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 04:47 PM
    Vision-language-action (VLA) models have emerged as a promising paradigm for autonomous driving.
    However, existing VLA models still suffer from a fundamental mismatch: driving actions require precise 3D geometric cues, while visual-language understanding and reasoning are largely conducted in a 2D semantic space.

Extractive summary: sentences quoted from the sources.