AION
Research paperMultimodal Models · Robotics & Embodied AI · Computer Vision2 sources · Oct 8, 2026

SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models

Existing spatial reasoning benchmarks mainly test spatial perception: reading off relations already visible in the input.

Key points

  • Yet real-world spatial intelligence demands predictive spatial reasoning: constructing a scene from observations, anticipating how an intervention changes it, and reasoning about the unseen outcome.
  • We introduce SpaceCast-Bench, the first benchmark to directly and diagnostically evaluate this capability.
  • Built around an observe-transform-infer framework, its 3,862 questions from 182 real-world scenes span 16 task types at three levels: static perception, local prediction, and global prediction, progressively requiring scene understanding, spatial state updating, and relational inference over unobserved outcomes.
  • Controlled analyses further reveal that bridge views are critical for integrating distributed observations, and that explicit 3D evidence benefits models more reliably than generated outcome images or videos.

Sources (2)

Extractive summary: sentences quoted from the sources.