AION
Research paperEfficiency & Inference · Computer Vision1 source · Oct 6, 2026

Later Is Better: Token Reduction for ViTs Under Distribution Shift

Training-free token reduction accelerates vision transformers by removing redundant tokens across layers, recovering most of the original accuracy at a fraction of the compute.

Key points

  • We show that this gap is governed by the reduction schedule, the depth profile of removal, usually left fixed as an implementation detail.
  • Concretely, we introduce a one-parameter late-concentrated power-law schedule that consistently improves out-of-distribution accuracy over flat at no extra inference cost.
  • On ImageNet-C with DeiT-S, the late schedule closes 83% of that gap at a 26% compute reduction (+1.17pp), and 99% of it at a lighter 7% reduction (+0.26pp).
  • Single-layer probes point to a mechanism: earlier reductions perturb features that pass through more remaining layers, front-loading reduction error in depth.

Sources (1)

  • [1]Later Is Better: Token Reduction for ViTs Under Distribution Shift
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 04:54 AM
    Training-free token reduction accelerates vision transformers by removing redundant tokens across layers, recovering most of the original accuracy at a fraction of the compute.
    We show that this gap is governed by the reduction schedule, the depth profile of removal, usually left fixed as an implementation detail.

Extractive summary: sentences quoted from the sources.