AION
Research paperRobotics & Embodied AI1 source · Oct 8, 2026

Unifying Policy Learning and State Prediction through Spatial Language Modeling

We introduce Spatial Language Modeling, which represents scene contours, goals, action targets, and future states with a shared vocabulary of discrete coordinates and semantic tokens.

Key points

  • Learning how actions change scene geometry can provide complementary supervision for goal-directed manipulation.
  • A task-specific grammar organizes these elements into spatial sequences, allowing one autoregressive Transformer to learn action generation and action-conditioned state prediction through a common next-token objective.
  • We train the model from scratch using random-play transition pretraining followed by joint action and state training on expert demonstrations.
  • During pretraining, recorded action coordinates condition subsequent state predictions and are excluded from the prediction loss.

Sources (1)

  • [1]Unifying Policy Learning and State Prediction through Spatial Language Modeling
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 03:43 PM
    We introduce Spatial Language Modeling, which represents scene contours, goals, action targets, and future states with a shared vocabulary of discrete coordinates and semantic tokens.
    Learning how actions change scene geometry can provide complementary supervision for goal-directed manipulation.

Extractive summary: sentences quoted from the sources.