AION
Research paperRobotics & Embodied AI1 source · Oct 8, 2026

Recompose and Refine Latent Reasoning Flows for Vision-Language-Action Models

Latent reasoning enables vision-language-action (VLA) models to transform multimodal observations into task-relevant internal states before generating continuous robot actions.

Key points

  • We present Reasoning and Flow Memory (FLOWMEM), a unified VLA model that turns successful latent computation into reusable reasoning experience.
  • Rather than appending a fixed retrieved context, FLOWMEM dynamically retrieves and recomposes compatible latent fragments as the embodied context evolves, forming a reasoning route that follows the temporal structure and progress of successful computation.
  • Experiments on RoboMME and LIBERO-Plus show that FLOWMEM attains 48.0% and 77.3% success, outperforming memory-free policies by 1.7 and 4.1 percentage points, respectively.
  • These results demonstrate the value of reusing successful latent computation for closed-loop VLA control.

Sources (1)

  • [1]Recompose and Refine Latent Reasoning Flows for Vision-Language-Action Models
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 03:00 PM
    Latent reasoning enables vision-language-action (VLA) models to transform multimodal observations into task-relevant internal states before generating continuous robot actions.
    We present Reasoning and Flow Memory (FLOWMEM), a unified VLA model that turns successful latent computation into reusable reasoning experience.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 8, 2026Rewiring Semantics, Dynamics, and Control: A Simple yet Effective Action-Centric Tri-Stream Transformer
  2. Oct 8, 2026Embodied Turing Machines: Stateful Code for Robot Recursive Self-Improvement
  3. Oct 7, 2026Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving
  4. Oct 7, 2026Do Vision-Language-Action Models Understand Instructions? A Mechanistic Interpretability Study on Language Grounding
  5. Oct 7, 2026Many Ways to Succeed: Diversity-Driven RL Fine-Tuning for VLA Generalization
  6. Oct 7, 2026Q-Learning with Scalar Adjoint Matching

Related