NTH

Decoding the Functional Roles of Register and High-Norm Patch Tokens in Vision Transformers

AuthorsNeel Varma, Andrew Rufail, Dipika Khullar, Vasu Sharma

October 10, 2026 2 min read
Watch on YouTube
The one-line take

This study shows that ViT register tokens carry crucial high-level semantics, while high-norm patch tokens mainly encode lower-level visual structure and contribute little to overall representations.

Key results

2.37%
High-norm outlier patch-token share

Top-norm layer-8 patch tokens selected for outlier analysis.

48.17%
Top-five register-feature ablation cosine-similarity drop

Representation stability change after ablating the five most active register-derived features.

0.31%
Top-five outlier-feature ablation cosine-similarity drop

Representation stability change after ablating the five most active outlier-derived features.

47.05%
Bottom-five register-feature ablation cosine-similarity drop

Shows that ablating less-active register features also strongly disrupts representation stability.

What the paper found

This study probes what register tokens and high-norm patch outliers do inside DINOv2, comparing standard and register-augmented DINOv2-small models at layer 8. The researchers trained sparse autoencoders on ImageNet-1k activations, then interpreted features with UMAP, CLIP-space comparisons, and patch descriptions generated by Gemma 3. Register-derived features aligned more with high-level visual semantics, while outlier features—selected from the top 2.37% of patch tokens by norm—were associated more with background structure and texture. Causal tests revealed a sharp functional difference: removing the five most active register-derived features reduced representation cosine similarity by 48.17%, compared with 0.31% for outlier-derived features. Removing the five least active register features still caused a 47.05% drop, suggesting that disruption depends less on which semantic features are removed than on exceeding the register system’s capacity. The findings support a division of labor in which registers help stabilize semantic representations while outlier tokens carry texture-heavy information. The authors caution that these analyses are limited to DINOv2 Small at layer 8, so whether the pattern holds across model sizes and layers remains untested.

Original abstract

Self-supervised Vision Transformers (ViTs), such as DINOv2, learn rich visual representations, but the functions of their internal tokens remain poorly understood. Recent architectures introduce dedicated register tokens to reduce high-norm out- lier patch tokens that emerge in background re- gions, yet the semantic and functional roles of both token types have not been fully established. In this paper, we analyze these roles by training sparse autoencoders (SAEs) on register-token and outlier-token activations in DINOv2. Using an automated interpretability pipeline, UMAP clus- tering, and CLIP-space cross-checks, we find that register-token features are more strongly associ- ated with high-level semantic concepts. Outlier- token features, by contrast, are more often associ- ated with lower-level structural, background, and texture-dominant patterns. Causal ablations fur- ther reveal a substantial functional asymmetry: disrupting top-activating register-derived features produces a 48.17% drop in representation cosine similarity, whereas disrupting outlier-derived fea- tures produces only a 0.31% drop. Together, our results provide evidence for token specialization in self-supervised ViTs.

Read the original paper

More in Transformers

Browse all 45 papers →
01Transformer

Length Generalization Needs Proper Regularization

Pavlo Vasylenko, Matthias Lindemann, André F. T. Martins, Marcos Treviso

The paper shows that carefully placed dropout can help language models generalize far beyond their training context length.

Read analysis
02Transformer

WavePrune: One period is often enough for RoPE

Guancheng Du, Luotian Huang, Shaowen Wang, Si Li, Kaifeng Lyu

WavePrune trims redundant RoPE rotations to improve long-context performance while making attention faster.

Read analysis