Bridging Compute- and Data-Optimal Pretraining
AuthorsTian Qin, Kimia Hamidieh, David Alvarez-Melis
Resources
This paper helps answer how to spend compute wisely when high-quality training data is scarce by modeling the diminishing value of repeated and paraphrased tokens.
Key results
Smallest OLMo3 model used to fit CD-scaling laws.
Largest model used in the primary experiments.
Size of the Dolma-3 pretraining corpus.
RMSE on log validation loss when fitting with 14M and 30M models and predicting larger models.
Approximate relative LOO-RMSE rejection of the constant-saturation baseline.
Model size at or above which paraphrasing is predicted to become ineffective for large data budgets.
What the paper found
Tian Qin, Kimia Hamidieh, and David Alvarez-Melis, from Harvard and MIT CSAIL, propose Compute-Data, or CD, scaling laws to extend Chinchilla-style compute-optimal scaling into the data-constrained regime. Their key variable is token effectiveness, η: a derived token from repetition or paraphrasing is converted into a fresh-token equivalent, with effective data defined as D plus ηD′. Using the OLMo3 architecture, the Dolma-3 150B corpus, and models from 14M to 600M parameters, they show that η is not constant: it decreases with model size, tokens-per-parameter ratio, and expansion amount, eventually saturating. The resulting power-law saturation model identifies three operational regimes—compute-bound, data-bound, and model-bound—and predicts a joint compute–data Pareto frontier rather than Chinchilla’s single allocation path. A fit using 14M and 30M models extrapolates to larger held-out models with RMSE 0.079 on log validation loss. Ablations show that a constant-saturation baseline has roughly 40% higher relative leave-one-out RMSE, confirming that dependence on model size and data availability is essential. Strategy recommendations are scale-dependent: paraphrasing is favored for models below 600M and limited data, while repetition becomes stronger as scale and tokens-per-parameter increase; extrapolation suggests four epochs are appropriate only around 3B-parameter models near Chinchilla data, and paraphrasing becomes ineffective at 7B or larger with data budgets of at least 4× Chinchilla.
Original abstract
Classical compute-optimal scaling laws assume an unbounded supply of fresh pretraining data, yet pretraining is increasingly entering a regime in which compute grows faster than the availability of high-quality data. We propose Compute-Data (CD) scaling laws, a unified framework that bridges compute-optimal scaling, where data scales freely with compute, and data-optimal scaling, where the corpus is fixed while compute can grow without bound. CD scaling extends classical scaling laws by introducing a token-effectiveness function, $η$, which quantifies the value of a derived token-produced, for example, through multi-epoch repetition or paraphrasing-relative to a fresh token, ranging from a perfect substitute to having no value. We fit $η$ for two data-expansion strategies, multi-epoch repetition and paraphrasing, across model sizes from 14M to 600M parameters using the Dolma-3 corpus. We find that token effectiveness is far from constant: it depends jointly on model size, the tokens-per-parameter ratio, and the amount of derived data, and it saturates as the corpus is expanded. The functional form of $η$ implies diminishing returns when substituting compute for data as either model size or data availability increases. It also partitions training into three operational regimes---compute-bound, data-bound, and model-bound---and shows that classical compute-optimal allocation is suboptimal across most practically relevant settings.
Read the original paperMore in Foundation Models
Browse all 47 papers →How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi
AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.
TabFM: A Zero-Shot Foundation Model for Tabular Data
Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das
TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.
When Do Biological Reasoning Models Use Their Biological Inputs?
Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik
The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.