NTH

Cracks in the Foundation: Seemingly Minor Architectural Choices Impact Long Context Extension

AuthorsAmanda Bertsch, Luca Soldaini, Matthew R. Gormley, Graham Neubig, Hannaneh Hajishirzi, Kyle Lo, Dirk Groeneveld

August 14, 2026 2 min read
Watch on YouTube
The one-line take

Small Transformer design choices can quietly make or break a model’s ability to handle very long contexts.

Key results

26
OlmPool model count

Comparable 7B models spanning four architectural choices

140B
Pretraining tokens

Tokens used before long-context extension

29.9
HELMET 32K minimum

Lowest score across OlmPool models

56.4
HELMET 32K maximum

Highest score across OlmPool models

0.67
Feature-count predictor R2

In-sample predictive power of counting detrimental features

170000
Training compute

GPU hours used for the 26-model study

What the paper found

This paper shows that long-context capability is determined partly by architectural decisions that appear minor in ordinary training. Using controlled experiments over Meta’s Llama 3, Qwen 3, and Olmo 3 design choices, the researchers varied normalization, grouped-query attention, sliding-window attention, and pretraining context length while holding data, tokenizer, and extension recipe fixed. The resulting OlmPool contains 26 comparable 7B models, each pretrained for 140B tokens and extended with 10B tokens of 64K-context data. On HELMET at 32K, scores range from 29.9 to 56.4, while RULER ranges from 44.7 to 67.7. Individually, most choices cause modest damage, but combinations compound: GQA plus sliding-window attention reduced HELMET by 9 points on average, and replacing Olmo 3’s QK normalization and post-sublayer normalization with prenorm improved it by 6 points. Standard short-context signals fail to expose these differences; counting the number of detrimental architectural features predicts long-context performance with an in-sample R2 of 0.67, compared with weak correlations from training loss and validation benchmarks. Mechanistically, QK norm produces higher-entropy, less concentrated attention and suppresses attention-sink behavior, while stronger attention sinks correlate with better long-context results at R2 0.38. The study required 170000 GPU hours and releases OlmPool checkpoints to support cheaper architecture research.

Original abstract

One might imagine that architectural variations within the dense transformer paradigm have a limited effect on accuracy. However, we demonstrate that this is not the case in the long context setting. Specifically, we show that a set of four minor architectural decisions --- all made by at least one of the Olmo, Llama, and Qwen dense model families --- have a compoundingly negative effect on long context extensibility. Any one of these choices alone has a minor impact on long context performance, but combining three or more can drop the performance downstream by up to 47%. Furthermore, these differences are not detectable from short-context loss or validation datasets. We show that much of the variation in long context ability across model families is driven by these architectural features and detectable from applying context extension early in pretraining. We demonstrate this with controlled ablations that hold data, tokenizer, and extension recipe fixed while varying normalization, GQA, pretraining context length, and sliding window attention. After over 170,000 GPU hours of training, we release the resulting set of models as OlmPool, a set of 26 comparable 7B models with checkpoints before and after long-context extension. This pool includes several architectures that outperform the Llama 3 architecture on long context extensibility. In an analysis of our ablation models, we identify patterns in attention sink behavior and attention distributions across context that are attributable to specific architectural differences.

Read the original paper

More in Transformers

Browse all 42 papers →
03Transformer

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Zehao Jin, Ruixuan Deng, Junran Wang

A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.

Read analysis