CARE: Certifying Acceleration for Vision-Language-Action Inference
Prior work accelerates VLA inference using techniques like action chunking and visual-token pruning, typically evaluating based on latency and average task success.
Key points
- While vision-language-action (VLA) models have advanced rapidly, running them at every control step remains expensive.
- We therefore define an acceleration-induced failure via paired rollouts from identical initial conditions, tracking when the reference succeeds but the accelerated policy fails.
- To manage this, we introduce CARE, an approach for certified accelerator selection.
- On four LIBERO suites with OpenVLA-OFT, CARE certifies 9.0--10.8times speedups while guaranteeing (at 95% confidence) that at least 85.8% of reference-solved episodes are preserved.
Sources (2)
- [1]CARE: Certifying Acceleration for Vision-Language-Action InferenceHugging Face Daily Papers · Oct 6, 12:00 AM
Prior work accelerates VLA inference using techniques like action chunking and visual-token pruning, typically evaluating based on latency and average task success.
While vision-language-action (VLA) models have advanced rapidly, running them at every control step remains expensive.
- [2]CARE: Certifying Acceleration for Vision-Language-Action InferencearXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 06:00 PM · same content
Extractive summary: sentences quoted from the sources.