NTH

Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs

AuthorsPavel Tikhonov, Anton Korznikov, Matvey Mikhalchuk, Nikita Dragunov, Temurbek Rahmatullaev, Polina Druzhinina, Anton Razzhigaev, Ivan Oseledets, Elena Tutubalina

September 28, 2026 2 min read
Watch on YouTube
The one-line take

This work argues that Transformers can blend two thoughts in one computation and proposes a way to separate them into coherent parallel continuations.

Key results

65%
Top-100 token recovery

Maximum reported recovery rate for a stream’s preferred next token in the mixed distribution.

1.86
Pythia-2.8B base KL divergence

Mean KL divergence between mixed-input predictions and the ideal average distribution before fine-tuning.

0.27
Pythia-2.8B tuned KL divergence

Mean KL divergence after lightweight self-distillation.

0.025%
Fine-tuning data fraction

The fine-tuning used less than this fraction of the original pretraining-data scale.

0.430
Llama-3.2-3B Joint Contrastive LAMBADA accuracy

Mean accuracy on superposed inputs using the proposed guided decoder.

What the paper found

This paper argues that ordinary decoder-only Transformers exhibit a measurable Superposition Linearity Hypothesis: when token embeddings from two unrelated text streams are averaged position by position, a single forward pass retains information from both, rather than producing semantic collapse. Across Pythia, Llama, Qwen, OLMo, and Gemma models evaluated on TinyStories and FineWeb, each stream’s preferred next token appears in the mixed distribution’s top-10 in roughly 30–40% of cases and reaches 65% recovery by top-100, far above a frequency-only baseline. Training analysis with Pythia checkpoints shows that this behavior is strongest at initialization and degrades during pretraining, suggesting an architectural bias rather than a learned capability; the residual stream’s final third remains especially linear. A lightweight self-distillation objective, using FineWeb pairs and less than 0.025% of the original pretraining-data scale, restores the property: for Pythia-2.8B, mean KL divergence from the ideal average distribution falls from 1.86 to 0.27. The paper also identifies a decoding obstacle: averaging logits behaves like a geometric mean, which can suppress tokens specific to either stream. Its Joint Contrastive decoder uses a smaller guide model to separate the streams, raising LAMBADA accuracy on superposed inputs to 0.430 for Llama-3.2-3B, versus 0.182 for pretrained mixing, though still below the 0.540 single-stream guide baseline. The experiments include OpenAI’s GPT-5.2 as a TinyStories judge, while the proposed approach points toward theoretically 2× inference throughput and reduced KV-cache use, subject to quality and context-length limitations.

Original abstract

While Large Language Models (LLMs) rely on highly non-linear components, in this work we demonstrate that they exhibit fundamental linearity: when inputs from distinct text streams are linearly combined, the model outputs a superposition of the individual next-token distributions. We term this the \textit{Superposition Linearity Hypothesis}. We provide evidence that superposition is an intrinsic property of the Transformer architecture rather than an emergent consequence of training; in fact, we observe that it tends to diminish as pretraining progresses. However, we demonstrate that linearity can be substantially restored through lightweight fine-tuning, significantly reducing the divergence between the predicted next-token distribution and the average of the individual next-token distributions. Finally, we introduce a guided decoding procedure that disentangles superposed outputs, enabling the simultaneous generation of two coherent continuations from a single forward pass.

Read the original paper

More in Transformers

Browse all 42 papers →
03Transformer

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Zehao Jin, Ruixuan Deng, Junran Wang

A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.

Read analysis