ResearchResearch paperRobotics & Embodied AI · Multimodal Models1 source · Oct 6, 2026

SpaTime: Streaming Vision-Language Models for Spatio-temporal Reasoning

We present SpaTime, a streaming VLM that fuses causal geometry tokens into the language model at every frame, using only the frames observed so far.

Key points

  • Embodied agents must reason about 3D space while the video is still arriving, answering questions as soon as they have observed enough of the scene.
  • VLMs that incorporate 3D geometric priors achieve strong spatial reasoning, but they operate offline, i.e., the full video must be available before they produce an answer.
  • Streaming VLMs process frames causally and decide for themselves when to respond, yet they lack explicit 3D representations.
  • To supervise when the model answers, we propose a response-time loss that maps per-frame response probabilities to a differentiable expected response time and penalizes the distance from the ground-truth frame.

Sources (1)

  • [1]SpaTime: Streaming Vision-Language Models for Spatio-temporal Reasoning
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 05:22 PM
    We present SpaTime, a streaming VLM that fuses causal geometry tokens into the language model at every frame, using only the frames observed so far.
    Embodied agents must reason about 3D space while the video is still arriving, answering questions as soon as they have observed enough of the scene.

Extractive summary: sentences quoted from the sources.

Related