Decoding the Functional Roles of Register and High-Norm Patch Tokens in Vision Transformers
AuthorsNeel Varma, Andrew Rufail, Dipika Khullar, Vasu Sharma
This study shows that ViT register tokens carry crucial high-level semantics, while high-norm patch tokens mainly encode lower-level visual structure and contribute little to overall representations.
Key results
Top-norm layer-8 patch tokens selected for outlier analysis.
Representation stability change after ablating the five most active register-derived features.
Representation stability change after ablating the five most active outlier-derived features.
Shows that ablating less-active register features also strongly disrupts representation stability.
What the paper found
This study probes what register tokens and high-norm patch outliers do inside DINOv2, comparing standard and register-augmented DINOv2-small models at layer 8. The researchers trained sparse autoencoders on ImageNet-1k activations, then interpreted features with UMAP, CLIP-space comparisons, and patch descriptions generated by Gemma 3. Register-derived features aligned more with high-level visual semantics, while outlier features—selected from the top 2.37% of patch tokens by norm—were associated more with background structure and texture. Causal tests revealed a sharp functional difference: removing the five most active register-derived features reduced representation cosine similarity by 48.17%, compared with 0.31% for outlier-derived features. Removing the five least active register features still caused a 47.05% drop, suggesting that disruption depends less on which semantic features are removed than on exceeding the register system’s capacity. The findings support a division of labor in which registers help stabilize semantic representations while outlier tokens carry texture-heavy information. The authors caution that these analyses are limited to DINOv2 Small at layer 8, so whether the pattern holds across model sizes and layers remains untested.
Original abstract
Self-supervised Vision Transformers (ViTs), such as DINOv2, learn rich visual representations, but the functions of their internal tokens remain poorly understood. Recent architectures introduce dedicated register tokens to reduce high-norm out- lier patch tokens that emerge in background re- gions, yet the semantic and functional roles of both token types have not been fully established. In this paper, we analyze these roles by training sparse autoencoders (SAEs) on register-token and outlier-token activations in DINOv2. Using an automated interpretability pipeline, UMAP clus- tering, and CLIP-space cross-checks, we find that register-token features are more strongly associ- ated with high-level semantic concepts. Outlier- token features, by contrast, are more often associ- ated with lower-level structural, background, and texture-dominant patterns. Causal ablations fur- ther reveal a substantial functional asymmetry: disrupting top-activating register-derived features produces a 48.17% drop in representation cosine similarity, whereas disrupting outlier-derived fea- tures produces only a 0.31% drop. Together, our results provide evidence for token specialization in self-supervised ViTs.
Read the original paperMore in Transformers
Browse all 45 papers →Length Generalization Needs Proper Regularization
Pavlo Vasylenko, Matthias Lindemann, André F. T. Martins, Marcos Treviso
The paper shows that carefully placed dropout can help language models generalize far beyond their training context length.
WavePrune: One period is often enough for RoPE
Guancheng Du, Luotian Huang, Shaowen Wang, Si Li, Kaifeng Lyu
WavePrune trims redundant RoPE rotations to improve long-context performance while making attention faster.
Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.