AION
Research paperImage, Video & 3D Generation1 source · Oct 6, 2026

S2Tok: Streaming 3D Gaussian Reconstruction with Persistent Spatial Tokens

We introduce S2Tok, a feed-forward framework that maintains a size-adaptive, persistent scene state from uncalibrated image streams.

Key points

  • Streaming 3D reconstruction requires more than a sequence of geometric predictions: it requires a persistent scene state that can incorporate new evidence and remain renderable as observations arrive.
  • A hierarchical decoder and Gaussian head convert the evolving state into non-pixel-aligned 3D Gaussians, enabling novel-view rendering without caching previous frames.
  • Experiments across four benchmarks demonstrate competitive streaming rendering quality with compact Gaussian representations.
  • These results support latent spatial tokens as a persistent computational state for online 3D reconstruction, combining learned scene updates with explicit Gaussian rendering.

Sources (1)

  • [1]S2Tok: Streaming 3D Gaussian Reconstruction with Persistent Spatial Tokens
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 06:40 PM
    We introduce S2Tok, a feed-forward framework that maintains a size-adaptive, persistent scene state from uncalibrated image streams.
    Streaming 3D reconstruction requires more than a sequence of geometric predictions: it requires a persistent scene state that can incorporate new evidence and remain renderable as observations arrive.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026huggingface/transformers v5.19.0: Release v5.19.0
  2. Oct 5, 2026LiquidAI/d1-omni-600M
  3. Oct 5, 2026MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers
  4. Sep 30, 2026Cloudflare/clef-flash
  5. Sep 29, 2026microsoft/AesCode-32B
  6. Aug 26, 2026vllm-project/vllm v0.28.0

Related