NTH

Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects

AuthorsPhu Gia Hoang, Anwoy Chatterjee, Tanmoy Chakraborty, Iryna Gurevych, Subhabrata Dutta

August 6, 2026 2 min read
Watch on YouTube
The one-line take

This study finds that sparse autoencoder features can matter causally without acting like simple, reusable steering directions, especially when they function as context-sensitive pointers.

Key results

65k
SAE dictionary width

Width used for ReLU, TopK, and Matryoshka Batch TopK SAEs on Gemma-2-2B.

0.40
TopK LSC-WC feature overlap

Intersection-over-union between pointer-like features selected from Literal Sequence Copying and Word Content.

18.4%
TopK Word Content ablation accuracy

Accuracy after ablating 9 selected features, compared with 93.6% ± 4.0% for matched random controls.

239
Pointer-like candidates

Total selected pointer-like features evaluated with FEGA.

3
Pointer-like directed rays

Only 3 pointer-like features formed stable directed-ray geometries.

36.3%
TopK value-like low-dimensional share

Share of geometry-eligible TopK RAVEL features assigned global, residual, or related low-dimensional categories.

What the paper found

This paper introduces Feature-Effect Geometry Analysis, or FEGA, an unsupervised method that evaluates SAE interventions by propagating feature ablations through the frozen model tail and analyzing their downstream logit-effect clouds rather than relying on local decoder directions. Using Google’s Gemma-2-2B with 65k-width post-layer-12 SAEs—ReLU, TopK, and Matryoshka Batch TopK—the authors compare value-like features identified by RAVEL city-country attribute editing with pointer-like features from Literal Sequence Copying, Word Content, PrOntoQA, and Token Translation. Pointer-like features recur across tasks, with TopK reaching an intersection-over-union of 0.40 between Literal Sequence Copying and Word Content, compared with a maximum of 0.13 for ReLU, and they are causally important: ablating 9 TopK Word Content features reduces accuracy to 18.4%, versus 93.6% ± 4.0% for matched random ablations. Yet their downstream effects are rarely stable steering directions. Among 239 pointer-like candidates, 159 have undefined geometry; of the 80 mapped cases, 74 are unresolved high-dimensional or diffuse, while only 3 form directed rays and 3 occupy global low-dimensional subspaces. Value-like features are more structured: low-dimensional categories cover 14.0% of eligible ReLU, 36.3% of TopK, and 30.3% of Matryoshka Batch TopK features, but directed rays remain rare. The central conclusion is that an SAE feature can encode a coherent concept or function and contribute causally without corresponding to a reusable one-dimensional logit-space control vector.

Original abstract

The wide-scale use of sparse autoencoders (SAEs) as interpretability tools is limited by inconsistent links between SAE features and model behavior. Features with clear activation descriptions may have weak or unexpected causal effects; steering can vary across prompts or oppose the intended direction; and activation-based feature selection can miss features that produce the desired output change. Prior work has studied feature geometry inside the model, where features are computed. We instead study the geometry of changes in model logits caused by feature interventions. We introduce Feature-Effect Geometry Analysis (FEGA), an unsupervised framework that removes the same active SAE feature across contexts and analyzes the resulting cloud of logit changes. Across SAE variants, consistent one-dimensional effects are rare: few features behave like reusable directions. To interpret this variation, we distinguish value-like features, tied to static information such as factual attributes, from pointer-like features, associated with context-dependent operations. Value-like features more often exhibit structured, low-dimensional effects, although these effects typically span several directions. Pointer-like features, by contrast, predominantly exhibit diffuse effects. Our results show that a feature can be interpretable and causally relevant without providing a stable direction for steering.

Read the original paper

More in Foundation Models

Browse all 47 papers →
01Foundation Model

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi

AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.

Read analysis
02Foundation Model

TabFM: A Zero-Shot Foundation Model for Tabular Data

Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das

TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.

Read analysis
03Foundation Model

When Do Biological Reasoning Models Use Their Biological Inputs?

Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik

The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.

Read analysis