Compact Robot Policies Need Fine-Grained Visual Representations
To test this, we build CoRP (Compressed Representation Policy), a deliberately compact policy (48.9M parameters, no vision-language model and no video-generative prior) that factorizes into a representation extractor and a flow-matching action generator.
Key points
- Multi-task manipulation policies differ in architecture, scale, and pretrained priors all at once, so published comparisons cannot attribute performance to any single component.
- We argue that most of it comes from the visual representation, and that parameter scale and generative priors are largely incidental.
- Compression matters as much: resampling each view to 48 tokens beats passing all patch tokens (97.0% vs 83.2%), and a variational information bottleneck over those tokens is worse than a hard token budget, cutting LIBERO-Goal from 95.8% to 33.0% by suppressing the instruction-dependent token selection the policy relies on.
- Therefore, we argue that a compact policy works when its representation is pretrained, task-adapted, and compressed.
Sources (1)
- [1]Compact Robot Policies Need Fine-Grained Visual RepresentationsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 11:36 AM
To test this, we build CoRP (Compressed Representation Policy), a deliberately compact policy (48.9M parameters, no vision-language model and no video-generative prior) that factorizes into a representation extractor and a flow-matching action generator.
Multi-task manipulation policies differ in architecture, scale, and pretrained priors all at once, so published comparisons cannot attribute performance to any single component.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 1, 2026nvidia/PixelDiT2-ImageNet
- Aug 26, 2026huggingface/transformers v5.16.0: Release: v5.16.0
- Aug 10, 2026vllm-project/vllm v0.27.0
- Aug 10, 2026huggingface/transformers v5.15.0: Release: v5.15.0
- Jun 29, 2026vllm-project/vllm v0.24.0
- Jun 15, 2026vllm-project/vllm v0.23.0