NTH

Bag of Dims: Training-Free Mechanistic Interpretability via Dimension-Level Sign Patterns

AuthorsVarun Reddy Nalagatla

July 15, 2026 3 min read
Watch on YouTube
The one-line take

The paper argues that transformer hidden states can be read directly in the default coordinate basis, with each dimension acting like a binary semantic register whose sign patterns reveal and even control concepts without training a probe.

Key results

0.006
Cross-dimensional mutual information

Pairwise dimension-sign mutual information remained below 0.006 bits.

0.99
Peak prototype AUC

Prototype detection reached AUC up to 0.99 across the evaluated language models.

100%
Unsupervised feature yield

Unsupervised discovery produced 1500 features with a 100% yield.

93%
Sign-only top-5 accuracy

Sign-only representations preserved up to 93% top-5 next-token accuracy.

92%
Training-free steering success

Closed-loop write-target steering achieved up to 92% concept-and-fluency success.

What the paper found

Varun Reddy Nalagatla of Amazon Web Services presents Bag of Dims, a training-free method arguing that transformer hidden states already expose semantic features in their ordinary coordinate system. Each dimension acts like a binary register: its sign carries content, while magnitude indicates strength, so concepts can be detected by counting sign agreements without sparse autoencoders, probes, labels, or learned rotations. Across Alibaba’s Qwen 3.5-4B and Qwen3-32B, Google’s Gemma 3-4B, Mistral 7B, Meta’s self-supervised DINOv2, Google’s supervised ViT-Base, and the Audio Spectrogram Transformer, cross-dimensional mutual information stayed below 0.006 bits, and MLPs added essentially no useful AUC over per-dimension reading. A single-token vocabulary cache detected 175 semantic categories at prototype AUC up to 0.99, while unsupervised discovery produced 1500 features with a 100% yield. Sign-only representations preserved 60–93% top-5 next-token accuracy, showing that magnitudes refine rather than define content. The paper separates detection from generation control: read dimensions can be sign-flipped to suppress concepts, whereas a distinct write target, computed as the sign of summed output-embedding rows, can be injected through attention output projections. Closed-loop steering induced concepts in fluent text with success rates of 62–92% across four language models, without training a steering vector. The results challenge the assumption, common in sparse-autoencoder work including Anthropic’s Claude interpretability research, that useful features require a learned rotation; the proposed bottleneck is instead cataloging what each layer-specific dimension encodes.

Original abstract

We show the standard basis of transformer hidden states already provides a training-free, architecture-general feature basis. Individual dimensions encode semantic content via their signs (+/-1) and confidence via their magnitudes, acting as independent binary registers; a feature is a subset of dimensions with a consistent sign pattern, read by counting sign agreements with no learned rotation. We validate this Bag of Dims framework across seven models spanning language (Qwen 3.5-4B, Gemma 3-4B, Mistral 7B, Qwen3-32B), vision (DINOv2, ViT-Base), and audio (AST). Signs alone carry predictive content: unit-magnitude sign patterns preserve 60-93% top-5 next-token accuracy through the LM head, and decoder-free Hamming scoring reaches 80-90% top-4096. From a single-token cache (one forward pass per token, no context, no labels), we detect 175 categories at AUC 0.97-0.99 by sign agreement; a trained probe adds only +0.018 AUC and converges to axis-aligned weights. These features are causally operative: they survive the K/V attention projections, trace to the FFN neuron coalitions that write them (random-weight controls never reproduce this), and flipping a feature's signs during the live forward pass suppresses its concept across four language models, magnitude-matched and concept-specific. Dimensions stay independent throughout (pairwise mutual information below 0.006 bits). The structure is not specific to language: the same per-dimension signs appear in self-supervised vision (DINOv2, 9/12 ImageNet superclasses), supervised vision (ViT-Base, 11/12), and audio (AST, 50/50 ESC-50 categories), so it reflects transformer training in general, not the language-modeling objective. The standard basis already suffices for feature reading at one forward pass, no optimization, no GPU-days. The open problem shifts from finding the right rotation to cataloging what each dimension encodes.

Read the original paper

More in AI Reasoning

Browse all 39 papers →
02Reasoning

On Language Drift during RLVR Post-Training

Michael Sullivan, Alexander Koller

RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.

Read analysis