NTH

Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention

AuthorsJing Huang, Daniel Wurgaft, Rachit Bansal, Laura Ruis, Naomi Saphra, David Alvarez-Melis, Andrew Kyle Lampinen, Christopher Potts, Ekdeep Singh Lubana

June 5, 2026 3 min read
Watch on YouTube
The one-line take

The paper explains why bigger models can learn rare, complex tasks that smaller ones miss: they have enough capacity to avoid overwriting those features while learning the common ones.

Key results

4B
OLMo model sizes

The paper validates the claims in OLMo pretraining using model sizes from 4M to 4B parameters, with the largest model being 4B.

7.58e-5
Non-task gradient cosine similarity, 1B

For the 1B model, the non-task gradient cosine similarity with the task direction is about 7.58 × 10^-5 ± 0.02, indicating near-orthogonality.

0.10
Non-task gradient cosine similarity, 20M

For the 20M model, the non-task gradient cosine similarity with the task direction is about 0.10 ± 0.09, showing much more interference than larger models.

What the paper found

This paper, from Stanford University, the Kempner Institute at Harvard, MIT, and Anthropic, argues that larger language models learn more not just because they are more expressive, but because scaling changes optimization dynamics on long-tailed task mixtures. Using a phenomenological reading of Chinchilla-style scaling, the authors distinguish tasks that smaller models can eventually recover with more data from tasks that require model scaling even with asymptotically unlimited data. In a synthetic mixture of orthogonal linear regression tasks, they show that width selects features by utility, defined as frequency times spectral magnitude, so small models spend capacity on high-frequency, low-complexity modes and fail on rare or complex ones. They then identify the key mechanism as reduced gradient interference: once a larger model has enough residual capacity to fit common tasks, gradients from those tasks shrink, leaving rare-task features less likely to be overwritten between sparse observations. In matched-frequency injection experiments, larger models retain rare-task signal across gaps, while narrow models exhibit an update-and-forget loop. The authors validate the same pattern in OLMo pretraining from 4M to 4B parameters on Dolma v1.7 by injecting two novel tasks, comparison and modular addition, at controlled frequencies. Only the larger OLMo models learn the lowest-frequency tasks, encode more task-specific features in their representations via distributed alignment search and Fourier analysis, and show near-orthogonal non-task gradients; for the 1B model, non-task gradient cosine similarity with the task direction is about 7.58×10^-5 ± 0.02, versus 0.10 ± 0.09 for the 20M model. The central claim is that memorization-like retention can be a prerequisite for learning rare, generalizable structure.

Original abstract

Larger models learn tasks smaller models do not. What drives this phenomenon? We develop a simple phenomenological argument that power-law scaling already suggests that a larger model will be able to learn a part of the data distribution that a smaller model fails to learn, even with infinite training data. To validate this claim and identify its causes, we study the effects of model scaling on a synthetic setup consisting of a mixture of tasks that show monotonic scaling curves. The results point to a data-induced competition over resources (neurons). Specifically, smaller models allocate their neurons to high frequency or low complexity tasks, and so they learn solutions that perform poorly on rare and complex tasks. Moreover, this happens even when solutions capable of expressing the desired task exist. We then assess how a larger model circumvents this data-centric bottleneck, finding that it traces to a reduced interference mechanism: larger models can allocate enough resources to common tasks that the gradient updates for those tasks become weak, which means that they do not overwrite rare-task features as they slowly accumulate. Finally, to further validate these claims, we pretrain OLMo models (4M to 4B parameters) on novel tasks of varying frequency and complexity. The results mirror those from our synthetic data experiments: only the larger OLMo models learn the infrequent and complex tasks, and these larger models embed more task features in their representations and show less gradient interference between tasks. Overall, we offer a data-centric account of why larger models learn tasks that smaller models fail to. This helps explain why larger models are better in practice, and it can inform practical questions concerning model sizing and training data mixtures.

Read the original paper

More in Foundation Models

Browse all 47 papers →
01Foundation Model

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi

AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.

Read analysis
02Foundation Model

TabFM: A Zero-Shot Foundation Model for Tabular Data

Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das

TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.

Read analysis
03Foundation Model

When Do Biological Reasoning Models Use Their Biological Inputs?

Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik

The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.

Read analysis