Dissecting Representation Structure in Vision Transformers: A Rigorous Architectural Study
Representation structure is crucial for understanding Vision Transformer (ViT) architectures and their generalization behavior.
ProofPaper ↗
Key points
- In this work, we conduct the first rigorous analysis of feature information across diverse architectural scales, empirically uncover the relationship between ViT representation and generalization behavior, and leverage these insights to guide efficient ViT design.
- Our contributions are fivefold: Across diverse architectural scales, 1) We identify feature collapse at initialization, which leads to redundancy, and propose a reduction scheme to mitigate this issue.
- 3) We show that feature in the token space provides a more faithful representation than those in embedding space.
- 4) We discover an unexpected finding: features produced by linear submodules within ViT layers are critical for the prediction of generalization performance.
Sources (1)
- [1]Dissecting Representation Structure in Vision Transformers: A Rigorous Architectural StudyarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 04:06 AM
Representation structure is crucial for understanding Vision Transformer (ViT) architectures and their generalization behavior.
In this work, we conduct the first rigorous analysis of feature information across diverse architectural scales, empirically uncover the relationship between ViT representation and generalization behavior, and leverage these insights to guide efficient ViT design.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 8, 2026MCL: Meta Convolution Layer
- Oct 8, 2026LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation
- Oct 7, 2026On the Necessity of Attention-FFN Split in Vision Transformers
- Oct 6, 2026huggingface/transformers v5.19.0: Release v5.19.0
- Oct 6, 2026EmbeddingGemma 2: an open, lightweight multimodal embedding model
- Oct 6, 2026Test-Time Adaptation of Quantized ViTs via Single-Pass Quantizer-Aligned Recalibration