NTH

Bridging Compute- and Data-Optimal Pretraining

AuthorsTian Qin, Kimia Hamidieh, David Alvarez-Melis

July 31, 2026 2 min read
Watch on YouTube
The one-line take

This paper helps answer how to spend compute wisely when high-quality training data is scarce by modeling the diminishing value of repeated and paraphrased tokens.

Key results

14M
Model-size sweep

Smallest OLMo3 model used to fit CD-scaling laws.

600M
Model-size sweep

Largest model used in the primary experiments.

150B
Fresh corpus

Size of the Dolma-3 pretraining corpus.

0.079
Held-out extrapolation RMSE

RMSE on log validation loss when fitting with 14M and 30M models and predicting larger models.

40%
Constant-saturation ablation

Approximate relative LOO-RMSE rejection of the constant-saturation baseline.

7B
Paraphrasing ineffectiveness threshold

Model size at or above which paraphrasing is predicted to become ineffective for large data budgets.

What the paper found

Tian Qin, Kimia Hamidieh, and David Alvarez-Melis, from Harvard and MIT CSAIL, propose Compute-Data, or CD, scaling laws to extend Chinchilla-style compute-optimal scaling into the data-constrained regime. Their key variable is token effectiveness, η: a derived token from repetition or paraphrasing is converted into a fresh-token equivalent, with effective data defined as D plus ηD′. Using the OLMo3 architecture, the Dolma-3 150B corpus, and models from 14M to 600M parameters, they show that η is not constant: it decreases with model size, tokens-per-parameter ratio, and expansion amount, eventually saturating. The resulting power-law saturation model identifies three operational regimes—compute-bound, data-bound, and model-bound—and predicts a joint compute–data Pareto frontier rather than Chinchilla’s single allocation path. A fit using 14M and 30M models extrapolates to larger held-out models with RMSE 0.079 on log validation loss. Ablations show that a constant-saturation baseline has roughly 40% higher relative leave-one-out RMSE, confirming that dependence on model size and data availability is essential. Strategy recommendations are scale-dependent: paraphrasing is favored for models below 600M and limited data, while repetition becomes stronger as scale and tokens-per-parameter increase; extrapolation suggests four epochs are appropriate only around 3B-parameter models near Chinchilla data, and paraphrasing becomes ineffective at 7B or larger with data budgets of at least 4× Chinchilla.

Original abstract

Classical compute-optimal scaling laws assume an unbounded supply of fresh pretraining data, yet pretraining is increasingly entering a regime in which compute grows faster than the availability of high-quality data. We propose Compute-Data (CD) scaling laws, a unified framework that bridges compute-optimal scaling, where data scales freely with compute, and data-optimal scaling, where the corpus is fixed while compute can grow without bound. CD scaling extends classical scaling laws by introducing a token-effectiveness function, $η$, which quantifies the value of a derived token-produced, for example, through multi-epoch repetition or paraphrasing-relative to a fresh token, ranging from a perfect substitute to having no value. We fit $η$ for two data-expansion strategies, multi-epoch repetition and paraphrasing, across model sizes from 14M to 600M parameters using the Dolma-3 corpus. We find that token effectiveness is far from constant: it depends jointly on model size, the tokens-per-parameter ratio, and the amount of derived data, and it saturates as the corpus is expanded. The functional form of $η$ implies diminishing returns when substituting compute for data as either model size or data availability increases. It also partitions training into three operational regimes---compute-bound, data-bound, and model-bound---and shows that classical compute-optimal allocation is suboptimal across most practically relevant settings.

Read the original paper

More in Foundation Models

Browse all 47 papers →
01Foundation Model

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi

AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.

Read analysis
02Foundation Model

TabFM: A Zero-Shot Foundation Model for Tabular Data

Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das

TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.

Read analysis
03Foundation Model

When Do Biological Reasoning Models Use Their Biological Inputs?

Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik

The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.

Read analysis