Bag of Dims: Training-Free Mechanistic Interpretability via Dimension-Level Sign Patterns
AuthorsVarun Reddy Nalagatla
Resources
The paper argues that transformer hidden states can be read directly in the default coordinate basis, with each dimension acting like a binary semantic register whose sign patterns reveal and even control concepts without training a probe.
Key results
Pairwise dimension-sign mutual information remained below 0.006 bits.
Prototype detection reached AUC up to 0.99 across the evaluated language models.
Unsupervised discovery produced 1500 features with a 100% yield.
Sign-only representations preserved up to 93% top-5 next-token accuracy.
Closed-loop write-target steering achieved up to 92% concept-and-fluency success.
What the paper found
Varun Reddy Nalagatla of Amazon Web Services presents Bag of Dims, a training-free method arguing that transformer hidden states already expose semantic features in their ordinary coordinate system. Each dimension acts like a binary register: its sign carries content, while magnitude indicates strength, so concepts can be detected by counting sign agreements without sparse autoencoders, probes, labels, or learned rotations. Across Alibaba’s Qwen 3.5-4B and Qwen3-32B, Google’s Gemma 3-4B, Mistral 7B, Meta’s self-supervised DINOv2, Google’s supervised ViT-Base, and the Audio Spectrogram Transformer, cross-dimensional mutual information stayed below 0.006 bits, and MLPs added essentially no useful AUC over per-dimension reading. A single-token vocabulary cache detected 175 semantic categories at prototype AUC up to 0.99, while unsupervised discovery produced 1500 features with a 100% yield. Sign-only representations preserved 60–93% top-5 next-token accuracy, showing that magnitudes refine rather than define content. The paper separates detection from generation control: read dimensions can be sign-flipped to suppress concepts, whereas a distinct write target, computed as the sign of summed output-embedding rows, can be injected through attention output projections. Closed-loop steering induced concepts in fluent text with success rates of 62–92% across four language models, without training a steering vector. The results challenge the assumption, common in sparse-autoencoder work including Anthropic’s Claude interpretability research, that useful features require a learned rotation; the proposed bottleneck is instead cataloging what each layer-specific dimension encodes.
Original abstract
We show the standard basis of transformer hidden states already provides a training-free, architecture-general feature basis. Individual dimensions encode semantic content via their signs (+/-1) and confidence via their magnitudes, acting as independent binary registers; a feature is a subset of dimensions with a consistent sign pattern, read by counting sign agreements with no learned rotation. We validate this Bag of Dims framework across seven models spanning language (Qwen 3.5-4B, Gemma 3-4B, Mistral 7B, Qwen3-32B), vision (DINOv2, ViT-Base), and audio (AST). Signs alone carry predictive content: unit-magnitude sign patterns preserve 60-93% top-5 next-token accuracy through the LM head, and decoder-free Hamming scoring reaches 80-90% top-4096. From a single-token cache (one forward pass per token, no context, no labels), we detect 175 categories at AUC 0.97-0.99 by sign agreement; a trained probe adds only +0.018 AUC and converges to axis-aligned weights. These features are causally operative: they survive the K/V attention projections, trace to the FFN neuron coalitions that write them (random-weight controls never reproduce this), and flipping a feature's signs during the live forward pass suppresses its concept across four language models, magnitude-matched and concept-specific. Dimensions stay independent throughout (pairwise mutual information below 0.006 bits). The structure is not specific to language: the same per-dimension signs appear in self-supervised vision (DINOv2, 9/12 ImageNet superclasses), supervised vision (ViT-Base, 11/12), and audio (AST, 50/50 ESC-50 categories), so it reflects transformer training in general, not the language-modeling objective. The standard basis already suffices for feature reading at one forward pass, no optimization, no GPU-days. The open problem shifts from finding the right rotation to cataloging what each dimension encodes.
Read the original paperMore in AI Reasoning
Browse all 39 papers →Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Principled Thoughts for Latent Recursive LLM Systems
Fahd Seddik, Fatemeh Fard
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.