Spatial Latent Reasoning for Embodied Reference Understanding
We propose Spatial Latent Reasoning (SLR), a framework that structures this supervision around an ordered sequence of geometric and visual states.
ProofPaper ↗
Key points
- Pointing-gesture visual grounding requires connecting hand geometry with the visual identity and extent of a referred object.
- A spatial ray state is supervised by fingertip position and pointing direction, followed by four states aligned with target-region features.
- To construct the visual targets, we introduce parity pooling, which applies polyphase grouping to average region tokens on four interleaved spatial supports.
- These results support task-structured supervision for continuous pointing grounding.
Sources (1)
- [1]Spatial Latent Reasoning for Embodied Reference UnderstandingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 04:20 AM
We propose Spatial Latent Reasoning (SLR), a framework that structures this supervision around an ordered sequence of geometric and visual states.
Pointing-gesture visual grounding requires connecting hand geometry with the visual identity and extent of a referred object.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 7, 2026The Persona Hierarchy Model: Understanding Contextual Generalization in Fine-Tuning LLMs
- Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
- Oct 6, 2026The Dichotomy Between Pattern Recognition and Step-by-Step Reasoning
- Oct 6, 2026HINTT Submission to the 2nd MLC-SLM Challenge: Comparing Cascaded and Unified Approaches to Diarization and ASR
- Oct 6, 2026Reading, Not Manipulating: Leveraging Router Logits for Multimodal Safety in MoE Vision-Language Models
- Jul 15, 2026huggingface/transformers v5.14.0: Release v5.14.0