Cracks in the Foundation: Seemingly Minor Architectural Choices Impact Long Context Extension
AuthorsAmanda Bertsch, Luca Soldaini, Matthew R. Gormley, Graham Neubig, Hannaneh Hajishirzi, Kyle Lo, Dirk Groeneveld
Resources
Small Transformer design choices can quietly make or break a model’s ability to handle very long contexts.
Key results
Comparable 7B models spanning four architectural choices
Tokens used before long-context extension
Lowest score across OlmPool models
Highest score across OlmPool models
In-sample predictive power of counting detrimental features
GPU hours used for the 26-model study
What the paper found
This paper shows that long-context capability is determined partly by architectural decisions that appear minor in ordinary training. Using controlled experiments over Meta’s Llama 3, Qwen 3, and Olmo 3 design choices, the researchers varied normalization, grouped-query attention, sliding-window attention, and pretraining context length while holding data, tokenizer, and extension recipe fixed. The resulting OlmPool contains 26 comparable 7B models, each pretrained for 140B tokens and extended with 10B tokens of 64K-context data. On HELMET at 32K, scores range from 29.9 to 56.4, while RULER ranges from 44.7 to 67.7. Individually, most choices cause modest damage, but combinations compound: GQA plus sliding-window attention reduced HELMET by 9 points on average, and replacing Olmo 3’s QK normalization and post-sublayer normalization with prenorm improved it by 6 points. Standard short-context signals fail to expose these differences; counting the number of detrimental architectural features predicts long-context performance with an in-sample R2 of 0.67, compared with weak correlations from training loss and validation benchmarks. Mechanistically, QK norm produces higher-entropy, less concentrated attention and suppresses attention-sink behavior, while stronger attention sinks correlate with better long-context results at R2 0.38. The study required 170000 GPU hours and releases OlmPool checkpoints to support cheaper architecture research.
Original abstract
One might imagine that architectural variations within the dense transformer paradigm have a limited effect on accuracy. However, we demonstrate that this is not the case in the long context setting. Specifically, we show that a set of four minor architectural decisions --- all made by at least one of the Olmo, Llama, and Qwen dense model families --- have a compoundingly negative effect on long context extensibility. Any one of these choices alone has a minor impact on long context performance, but combining three or more can drop the performance downstream by up to 47%. Furthermore, these differences are not detectable from short-context loss or validation datasets. We show that much of the variation in long context ability across model families is driven by these architectural features and detectable from applying context extension early in pretraining. We demonstrate this with controlled ablations that hold data, tokenizer, and extension recipe fixed while varying normalization, GQA, pretraining context length, and sliding window attention. After over 170,000 GPU hours of training, we release the resulting set of models as OlmPool, a set of 26 comparable 7B models with checkpoints before and after long-context extension. This pool includes several architectures that outperform the Llama 3 architecture on long context extensibility. In an analysis of our ablation models, we identify patterns in attention sink behavior and attention distributions across context that are attributable to specific architectural differences.
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.