NTH

Unified Neural Scaling Laws

AuthorsEthan Caballero, Priyank Jaini, David Krueger, Irina Rish

June 5, 2026 2 min read
Watch on YouTube
The one-line take

This paper proposes a single scaling-law formula that better predicts how neural networks behave as model size, data, compute, and other factors all change at once.

Key results

60.87%
Image Classification Best Rate

UNSL is the best extrapolation model on 60.87% of downstream image classification tasks, beating the next best form at 21.74%.

88.89%
Language Best Rate

UNSL is the best extrapolation model on 88.89% of language tasks, versus 11.11% for the next best form.

2.54e-1
ImageNet RMSLE DC

For ImageNet bivariate extrapolation, DC's RMSLE is reported as 2.54e-1.

8.57e-3
ImageNet RMSLE UNSL

For ImageNet bivariate extrapolation, UNSL's RMSLE is reported as 8.57e-3.

6.24e-2
Language RMSLE DC

For language trivariate extrapolation, DC's RMSLE is reported as 6.24e-2.

7.82e-3
Language RMSLE UNSL

For language trivariate extrapolation, UNSL's RMSLE is reported as 7.82e-3.

What the paper found

Unified Neural Scaling Laws, from Mila researchers Ethan Caballero, David Krueger, and Irina Rish with Priyank Jaini of Google DeepMind, proposes a single multivariate curve family that models deep-network performance as model size, dataset size, training steps, inference steps, batch size, and hyperparameters vary together. The key novelty is a nested functional form built from multivariate broken neural scaling laws with smooth “hyperbreaks,” plus explicit terms for bottlenecks, irreducible performance limits, overfitting, and nonmonotonic hyperparameter effects such as learning rate and initialization scale. Unlike earlier scaling laws such as the Kaplan and Hoffmann Chinchilla-style formulas, or Muennighoff et al.’s data-constrained language model law, UNSL can represent transitions that reverse direction and does so with the same asymptotic expressivity as simpler baselines. On held-out extrapolation, it is substantially more accurate: in downstream image classification it is the best model on 60.87% of tasks, versus 21.74% for the next best form, and in language it wins on 88.89% of tasks. Reported RMSLE drops include ImageNet bivariate extrapolation from 2.54e-1 for DC to 8.57e-3 for UNSL, and language bivariate extrapolation from 6.24e-2 for DC to 7.82e-3 for UNSL. The paper also shows the law extrapolating reinforcement learning, test-time chain-of-thought scaling, width-depth tradeoffs, and sparse parity, while revealing a practical limit: accurate forecasting past a regime change requires training data near the relevant hyperbreak.

Original abstract

We present a functional form (that we refer to as a Unified Neural Scaling Law (UNSL)) that accurately models and extrapolates the scaling behaviors of deep neural networks as multiple dimensions all vary simultaneously (i.e. how the evaluation metric of interest varies as one simultaneously varies the number of model parameters, training dataset size, number of training steps, number of inference steps, amount of compute, and various hyperparameters) for various architectures and for each of various tasks within a varied set of upstream and downstream tasks. This set includes large-scale vision, language, math, and reinforcement learning. When compared to other functional forms for neural scaling, this functional form yields extrapolations of scaling behavior that are considerably more accurate on this set.

Read the original paper

More in Foundation Models

Browse all 47 papers →
01Foundation Model

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi

AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.

Read analysis
02Foundation Model

TabFM: A Zero-Shot Foundation Model for Tabular Data

Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das

TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.

Read analysis
03Foundation Model

When Do Biological Reasoning Models Use Their Biological Inputs?

Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik

The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.

Read analysis