ResearchResearch paperRobotics & Embodied AI · Reinforcement Learning · Large Language Models1 source · Oct 6, 2026

CAP: Codebook-Aligned Prediction for Tokenized Robot Policies

Action tokenization converts continuous robot actions into discrete symbols that can be modeled autoregressively.

Key points

  • However, existing tokenizer-based policies typically ignore the tokenizer's learned latent code structure: after tokenization, the policy treats tokens as unrelated class indices and learns a new classifier from scratch.
  • We introduce Codebook-Aligned Prediction (CAP), a method that directly reuses the tokenizer's code vectors as policy class prototypes while leaving the tokenizer and policy backbone otherwise unchanged.
  • Across four quantizer families, three simulation benchmarks, and two real-robot tasks, CAP consistently improves task success over standard token classification heads while holding the tokenizer (and therefore its reconstruction quality) fixed.
  • These results suggest that action tokenizers learn useful action-aware latent structure beyond discrete targets that should be preserved when training downstream policies.

Sources (1)

  • [1]CAP: Codebook-Aligned Prediction for Tokenized Robot Policies
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 10:28 PM
    Action tokenization converts continuous robot actions into discrete symbols that can be modeled autoregressively.
    However, existing tokenizer-based policies typically ignore the tokenizer's learned latent code structure: after tokenization, the policy treats tokens as unrelated class indices and learns a new classifier from scratch.

Extractive summary: sentences quoted from the sources.

Related