SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models
Existing spatial reasoning benchmarks mainly test spatial perception: reading off relations already visible in the input.
Key points
- Yet real-world spatial intelligence demands predictive spatial reasoning: constructing a scene from observations, anticipating how an intervention changes it, and reasoning about the unseen outcome.
- We introduce SpaceCast-Bench, the first benchmark to directly and diagnostically evaluate this capability.
- Built around an observe-transform-infer framework, its 3,862 questions from 182 real-world scenes span 16 task types at three levels: static perception, local prediction, and global prediction, progressively requiring scene understanding, spatial state updating, and relational inference over unobserved outcomes.
- Controlled analyses further reveal that bridge views are critical for integrating distributed observations, and that explicit 3D evidence benefits models more reliably than generated outcome images or videos.
Sources (2)
- [1]SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language ModelsHugging Face Daily Papers · Oct 8, 12:00 AM
Existing spatial reasoning benchmarks mainly test spatial perception: reading off relations already visible in the input.
Yet real-world spatial intelligence demands predictive spatial reasoning: constructing a scene from observations, anticipating how an intervention changes it, and reasoning about the unseen outcome.
- [2]SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language ModelsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:44 PM · same content
Extractive summary: sentences quoted from the sources.