CAP: Codebook-Aligned Prediction for Tokenized Robot Policies
Action tokenization converts continuous robot actions into discrete symbols that can be modeled autoregressively.
ProofPaper ↗
Key points
- However, existing tokenizer-based policies typically ignore the tokenizer's learned latent code structure: after tokenization, the policy treats tokens as unrelated class indices and learns a new classifier from scratch.
- We introduce Codebook-Aligned Prediction (CAP), a method that directly reuses the tokenizer's code vectors as policy class prototypes while leaving the tokenizer and policy backbone otherwise unchanged.
- Across four quantizer families, three simulation benchmarks, and two real-robot tasks, CAP consistently improves task success over standard token classification heads while holding the tokenizer (and therefore its reconstruction quality) fixed.
- These results suggest that action tokenizers learn useful action-aware latent structure beyond discrete targets that should be preserved when training downstream policies.
Sources (1)
- [1]CAP: Codebook-Aligned Prediction for Tokenized Robot PoliciesarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 10:28 PM
Action tokenization converts continuous robot actions into discrete symbols that can be modeled autoregressively.
However, existing tokenizer-based policies typically ignore the tokenizer's learned latent code structure: after tokenization, the policy treats tokens as unrelated class indices and learns a new classifier from scratch.
Extractive summary: sentences quoted from the sources.