One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts
In this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combination of a small shared expert bank.
Key points
- A continuous normalized-depth coordinate programs this mixture, defining a resampleable trajectory through FFN parameter space.
- We evaluate this design in two regimes: supervised ImageNet-1k training and distillation from a DINOv2 teacher.
- Elastic-depth training allows one checkpoint (trained model) to operate at multiple tested depths by resampling the same normalized coordinate interval.
- For fixed-depth deployment, the recurrent block can be materialized as a conventional dense graph, removing online routing and merging without changing the one-FFN-per-depth compute but expanding deployment storage.
Sources (2)
- [1]One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed ExpertsHugging Face Daily Papers · Oct 8, 12:00 AM
In this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combination of a small shared expert bank.
A continuous normalized-depth coordinate programs this mixture, defining a resampleable trajectory through FFN parameter space.
- [2]One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed ExpertsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:58 PM · same content
Extractive summary: sentences quoted from the sources.