Unifying Policy Learning and State Prediction through Spatial Language Modeling
We introduce Spatial Language Modeling, which represents scene contours, goals, action targets, and future states with a shared vocabulary of discrete coordinates and semantic tokens.
Key points
- Learning how actions change scene geometry can provide complementary supervision for goal-directed manipulation.
- A task-specific grammar organizes these elements into spatial sequences, allowing one autoregressive Transformer to learn action generation and action-conditioned state prediction through a common next-token objective.
- We train the model from scratch using random-play transition pretraining followed by joint action and state training on expert demonstrations.
- During pretraining, recorded action coordinates condition subsequent state predictions and are excluded from the prediction loss.
Sources (1)
- [1]Unifying Policy Learning and State Prediction through Spatial Language ModelingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 03:43 PM
We introduce Spatial Language Modeling, which represents scene contours, goals, action targets, and future states with a shared vocabulary of discrete coordinates and semantic tokens.
Learning how actions change scene geometry can provide complementary supervision for goal-directed manipulation.
Extractive summary: sentences quoted from the sources.