Later Is Better: Token Reduction for ViTs Under Distribution Shift
Training-free token reduction accelerates vision transformers by removing redundant tokens across layers, recovering most of the original accuracy at a fraction of the compute.
Key points
- We show that this gap is governed by the reduction schedule, the depth profile of removal, usually left fixed as an implementation detail.
- Concretely, we introduce a one-parameter late-concentrated power-law schedule that consistently improves out-of-distribution accuracy over flat at no extra inference cost.
- On ImageNet-C with DeiT-S, the late schedule closes 83% of that gap at a 26% compute reduction (+1.17pp), and 99% of it at a lighter 7% reduction (+0.26pp).
- Single-layer probes point to a mechanism: earlier reductions perturb features that pass through more remaining layers, front-loading reduction error in depth.
Sources (1)
- [1]Later Is Better: Token Reduction for ViTs Under Distribution ShiftarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 04:54 AM
Training-free token reduction accelerates vision transformers by removing redundant tokens across layers, recovering most of the original accuracy at a fraction of the compute.
We show that this gap is governed by the reduction schedule, the depth profile of removal, usually left fixed as an implementation detail.
Extractive summary: sentences quoted from the sources.