Recompose and Refine Latent Reasoning Flows for Vision-Language-Action Models
Latent reasoning enables vision-language-action (VLA) models to transform multimodal observations into task-relevant internal states before generating continuous robot actions.
Key points
- We present Reasoning and Flow Memory (FLOWMEM), a unified VLA model that turns successful latent computation into reusable reasoning experience.
- Rather than appending a fixed retrieved context, FLOWMEM dynamically retrieves and recomposes compatible latent fragments as the embodied context evolves, forming a reasoning route that follows the temporal structure and progress of successful computation.
- Experiments on RoboMME and LIBERO-Plus show that FLOWMEM attains 48.0% and 77.3% success, outperforming memory-free policies by 1.7 and 4.1 percentage points, respectively.
- These results demonstrate the value of reusing successful latent computation for closed-loop VLA control.
Sources (1)
- [1]Recompose and Refine Latent Reasoning Flows for Vision-Language-Action ModelsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 03:00 PM
Latent reasoning enables vision-language-action (VLA) models to transform multimodal observations into task-relevant internal states before generating continuous robot actions.
We present Reasoning and Flow Memory (FLOWMEM), a unified VLA model that turns successful latent computation into reusable reasoning experience.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 8, 2026Rewiring Semantics, Dynamics, and Control: A Simple yet Effective Action-Centric Tri-Stream Transformer
- Oct 8, 2026Embodied Turing Machines: Stateful Code for Robot Recursive Self-Improvement
- Oct 7, 2026Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving
- Oct 7, 2026Do Vision-Language-Action Models Understand Instructions? A Mechanistic Interpretability Study on Language Grounding
- Oct 7, 2026Many Ways to Succeed: Diversity-Driven RL Fine-Tuning for VLA Generalization
- Oct 7, 2026Q-Learning with Scalar Adjoint Matching