Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs
AuthorsPavel Tikhonov, Anton Korznikov, Matvey Mikhalchuk, Nikita Dragunov, Temurbek Rahmatullaev, Polina Druzhinina, Anton Razzhigaev, Ivan Oseledets, Elena Tutubalina
Resources
This work argues that Transformers can blend two thoughts in one computation and proposes a way to separate them into coherent parallel continuations.
Key results
Maximum reported recovery rate for a stream’s preferred next token in the mixed distribution.
Mean KL divergence between mixed-input predictions and the ideal average distribution before fine-tuning.
Mean KL divergence after lightweight self-distillation.
The fine-tuning used less than this fraction of the original pretraining-data scale.
Mean accuracy on superposed inputs using the proposed guided decoder.
What the paper found
This paper argues that ordinary decoder-only Transformers exhibit a measurable Superposition Linearity Hypothesis: when token embeddings from two unrelated text streams are averaged position by position, a single forward pass retains information from both, rather than producing semantic collapse. Across Pythia, Llama, Qwen, OLMo, and Gemma models evaluated on TinyStories and FineWeb, each stream’s preferred next token appears in the mixed distribution’s top-10 in roughly 30–40% of cases and reaches 65% recovery by top-100, far above a frequency-only baseline. Training analysis with Pythia checkpoints shows that this behavior is strongest at initialization and degrades during pretraining, suggesting an architectural bias rather than a learned capability; the residual stream’s final third remains especially linear. A lightweight self-distillation objective, using FineWeb pairs and less than 0.025% of the original pretraining-data scale, restores the property: for Pythia-2.8B, mean KL divergence from the ideal average distribution falls from 1.86 to 0.27. The paper also identifies a decoding obstacle: averaging logits behaves like a geometric mean, which can suppress tokens specific to either stream. Its Joint Contrastive decoder uses a smaller guide model to separate the streams, raising LAMBADA accuracy on superposed inputs to 0.430 for Llama-3.2-3B, versus 0.182 for pretrained mixing, though still below the 0.540 single-stream guide baseline. The experiments include OpenAI’s GPT-5.2 as a TinyStories judge, while the proposed approach points toward theoretically 2× inference throughput and reduced KV-cache use, subject to quality and context-length limitations.
Original abstract
While Large Language Models (LLMs) rely on highly non-linear components, in this work we demonstrate that they exhibit fundamental linearity: when inputs from distinct text streams are linearly combined, the model outputs a superposition of the individual next-token distributions. We term this the \textit{Superposition Linearity Hypothesis}. We provide evidence that superposition is an intrinsic property of the Transformer architecture rather than an emergent consequence of training; in fact, we observe that it tends to diminish as pretraining progresses. However, we demonstrate that linearity can be substantially restored through lightweight fine-tuning, significantly reducing the divergence between the predicted next-token distribution and the average of the individual next-token distributions. Finally, we introduce a guided decoding procedure that disentangles superposed outputs, enabling the simultaneous generation of two coherent continuations from a single forward pass.
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.