S2Tok: Streaming 3D Gaussian Reconstruction with Persistent Spatial Tokens
We introduce S2Tok, a feed-forward framework that maintains a size-adaptive, persistent scene state from uncalibrated image streams.
Key points
- Streaming 3D reconstruction requires more than a sequence of geometric predictions: it requires a persistent scene state that can incorporate new evidence and remain renderable as observations arrive.
- A hierarchical decoder and Gaussian head convert the evolving state into non-pixel-aligned 3D Gaussians, enabling novel-view rendering without caching previous frames.
- Experiments across four benchmarks demonstrate competitive streaming rendering quality with compact Gaussian representations.
- These results support latent spatial tokens as a persistent computational state for online 3D reconstruction, combining learned scene updates with explicit Gaussian rendering.
Sources (1)
- [1]S2Tok: Streaming 3D Gaussian Reconstruction with Persistent Spatial TokensarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 06:40 PM
We introduce S2Tok, a feed-forward framework that maintains a size-adaptive, persistent scene state from uncalibrated image streams.
Streaming 3D reconstruction requires more than a sequence of geometric predictions: it requires a persistent scene state that can incorporate new evidence and remain renderable as observations arrive.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026huggingface/transformers v5.19.0: Release v5.19.0
- Oct 5, 2026LiquidAI/d1-omni-600M
- Oct 5, 2026MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers
- Sep 30, 2026Cloudflare/clef-flash
- Sep 29, 2026microsoft/AesCode-32B
- Aug 26, 2026vllm-project/vllm v0.28.0