The Linear Representation Hypothesis Needs a Group Action
AuthorsLouie Hong Yao, Yuhao Li, Shengchao Liu
Resources
This paper argues that claims about linear representations only become meaningful once we specify which transformations leave a representation essentially unchanged.
Key results
The hierarchy contains Giso, Gsim, and Gaff.
Steering displacements and probe weights occupy V and V*.
What the paper found
This paper argues that the Linear Representation Hypothesis is not a single claim: it is a family of claims determined by what transformations count as equivalent representations. Its proposed formal specification has four parts: an equivalence group G, an object space M, an equivariant extraction procedure F, and an invariant predicate P. The framework distinguishes three nested affine symmetry groups—Giso, Gsim, and Gaff—which respectively preserve fixed Euclidean geometry, angles up to scale, and only affine structure. This distinction explains why a steering displacement and a linear-probe weight cannot automatically be treated as the same direction: they live in two distinct spaces, the activation space V and its dual V*, with transformations v to Av and w to A−⊤w. For vector spaces of dimension at least 2, no nonzero GL(V)-equivariant map identifies these objects without added metric structure. The paper also introduces an architectural floor: function-preserving reparameterizations constrain the minimum symmetry a model-level claim must respect. In transformer attention, query-key gauge transformations can preserve the model function while changing cosine similarities, norms, PCA directions, and even KeyDiff eviction decisions at query or key reading points; RoPE reduces but does not eliminate this anisotropic symmetry. The audit applies to difference-in-means steering, Contrastive Activation Addition, sparse autoencoders, and analyses of Llama 2 and Claude 3 Sonnet, showing that pipelines must be checked as complete compositions rather than judged stage by stage. The central recommendation is to report the object, group action, procedure, predicate, and reading point explicitly.
Original abstract
To make claims about representations that generalize beyond a particular trained model, we need to specify when two representations should count as equivalent. The Linear Representation Hypothesis is often discussed without making this equivalence explicit. Different notions of equivalence preserve different structures, so metrics, probes, and interventions that appear to study the same representation may in fact correspond to different hypotheses. We therefore argue that the Linear Representation Hypothesis is not one hypothesis but a family of claims distinguished by representation equivalence. We formalize this idea using group actions, specifying the representation object, the procedure that produces it, and the property ultimately asserted, while accounting for equivalences imposed by the model architecture. This framework clarifies how assumptions can change across metrics, reading points, and analysis stages, and we use it to audit common representation quantities and recent interpretability analyses.
Read the original paperMore in Neural Networks
Browse all 22 papers →End-to-End Hard-Label Cryptanalytic Model Extraction Using Efficient Sign Recovery
Akira Ito, Takayuki Miura, Yosuke Todo
A new query-efficient technique makes it possible to steal the parameters of small black-box neural networks using only their predicted labels.
Retrieving Individual Stems from Music Mixtures with Slot Embeddings
David Braun, Junyi Fan, Pranay Manocha, Donald S. Williamson, Adam Finkelstein
Stembed lets music producers search for individual instrument sounds hidden inside a full song by representing the mixture as multiple searchable stem-like embeddings.
HypLTSF: A Hyperbolic Geometric View of Multi-Scale Hierarchies for Long-Term Time Series Forecasting
Namwoo Kim, Hyungryul Baik, Yoonjin Yoon
HypLTSF maps multi-scale time-series patterns into hyperbolic space so that fine-to-coarse temporal hierarchies become explicit and useful for long-term forecasting.