Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
AuthorsZehao Jin, Ruixuan Deng, Junran Wang
AffiliationsGeorgia Institute of Technology
Resources
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.
Key results
Median number of reference-chain lines reliably followed across 13 standard models.
Trainable parameters in the rank-8 single-layer adapter.
Exact accuracy after LoRA adaptation, compared with 15.5% for the frozen model.
Lower-bound chain reach after applying a longer-trained LoRA through eight loops.
Exact-match improvement from an early task-specific LoRA.
What the paper found
This paper argues that pretrained transformers often stop reference-following long before their available depth is exhausted. Across 13 standard models, including Qwen3, Llama, OLMo, Gemma, and the 292B-MoE DeepSeek-V4-Flash, reliable reach is only 1.4 to 3.6 assignment lines, with a median of 2.2, and extra recurrent loops in Ouro and Huginn add little. The proposed fix is a rank-8 residual-stream LoRA inserted at a model-specific early or middle layer while freezing every original weight. In Qwen3-8B, the adapter trains only 65537 parameters and raises exact accuracy on 24-line reference chains from 15.5% to 99%; a longer-trained version reaches 50 lines in one pass. In Ouro-1.4B, applying the same kind of adapter through recurrent loops extends reach to 60 lines after four loops and at least 160 after eight. Mechanistic tracing suggests the LoRA does not perform token-to-token communication itself; instead, it activates a relay in existing attention and MLP layers, where program lines progressively carry chain identity and ancestor names through a short middle-layer window. Parent-line attention is necessary, while removing it after the relay has formed has little effect. The placement is sharply constrained: moving the Qwen3-8B intervention from layer 20 to layer 21 reduces reach from 20.5 to 5.2 lines. On the multi-hop MuSiQue benchmark, early task-specific LoRAs improve exact-match performance by 11.4 points on Qwen3-8B, 9.4 on OLMo-3-7B, and 17.9 on Llama-3.1-8B, showing that the intervention exposes computation already latent in frozen models rather than simply adding capacity.
Original abstract
Pretrained transformers use little of their depth to follow references in context. Thirteen base models reliably follow only 1.4-3.6 lines, and extra pretrained loops add little. A task-trained rank-8 LoRA at one early layer extends this computation with all model weights frozen. Qwen3-8B improves from 15.5% to 99% exact accuracy on 24-line chains; a longer-trained LoRA reaches 50 lines. Ouro-1.4B reaches 60 lines after four loops and at least 160 after eight. The LoRA starts a relay: program lines pass on their chain identity through a short range of middle layers. Frozen heads read progressively further up the chain, and removing parent-line attention stops the relay. A frozen-model measurement locates the last useful intervention layer within tolerance in three of four held-out models. Task-specific LoRAs also improve MuSiQue. Default answers therefore understate the computation accessible through a tiny edit. Code and an interactive demo are available at https://lunamos.github.io/stop-thinking-too-early/
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Nonparametric In-Context Learning under Growing Geometric Complexity: Minimax Optimality and Local Geometry-Adaptivity of Transformers
Jaehee Seo, Jisu Kim
This work shows, in theory, how transformers can adapt to data living on locally different geometric structures and still achieve statistically optimal in-context prediction.