NTH

Small-Scale Experiments: Are We There Yet?

AuthorsNicholas Lourie, Kyunghyun Cho, Karen Ullrich, Sanae Lotfi

August 16, 2026 3 min read
Watch on YouTube
The one-line take

Small models can predict large-model behavior after all—but only when their hyperparameters are tuned carefully enough to reveal the hidden scaling laws.

Key results

4M
Smallest scaling-law scale

Scaling laws emerge at models with 4M effective parameters when hyperparameters are thoroughly tuned.

256
Accurate tuning sweep

The scaling law becomes accurate only after evaluating 256 configurations.

98%
Learning-rate decay gain

Learning-rate decay reduced held-out test mean-squared error by 98%.

1
Effective hyperparameter dimension

The noisy quadratic analysis shows the effective hyperparameter count falls to 1 as models scale.

100B
FineWeb-Edu training subset

Experiments use a 100B-token subset of FineWeb-Edu.

What the paper found

This paper argues that scaling laws do exist far below production-model scale, but small models are difficult to tune, so the laws remain hidden unless researchers reach the fully optimized hyperparameter frontier. Experiments with a Llama decoder in Meta Lingua, using OpenAI’s p50k_base tokenizer, the 100B-token FineWeb-Edu subset, and NVIDIA A100 GPUs, span 4M to 268M effective parameters. Random-search studies show that reliable scaling requires far more tuning than typical practice: the relationship is absent with only 4 or 16 configurations, becomes visible at 64, and is accurate at 256 configurations. Among methodological choices, learning-rate decay reduced held-out test mean-squared error by 98%, while parameter-count conventions and tied scaling exponents mattered much less. The proposed noisy quadratic limit reveals that hyperparameter loss surfaces become lower-dimensional as models grow, with the effective hyperparameter count falling to 1, explaining why large models are easier to tune. The authors combine this diagnostic with scaling laws and perplexity-capability correspondence, validating conclusions on AI2 ARC, BoolQ, MMLU, OpenBookQA, and PIQA while warning that extrapolation eventually becomes statistically unreliable. In a transformer normalization case study, small-scale experiments recover the established result that pre-normalization scales better than post-normalization: post-norm remains more hyperparameter-sensitive, whereas pre-norm yields a cleaner scaling law and stronger compute efficiency near observed data. The practical recommendation is to explore hundreds of configurations at small scale, verify diagnostics, and transfer the resulting design upward rather than relying on naïve extrapolation.

Original abstract

Scaling laws promised cost-effective experiments; six years later, they have yet to fully deliver. Instead, researchers have found them unreliable at small scales (starting at 4M parameters) and concluded that sizable models cannot be avoided. We show this is not the case: the confounding factor is hyperparameters. Small models are highly sensitive, but hyperparameter sensitivity fades with scale. This small-scale sensitivity makes scaling laws easy to miss because they only emerge on the fully tuned frontier, and reaching that frontier requires an extensive search far beyond what most ever run. By ablating the basic scaling law recipe, we show well-tuned hyperparameters matter more than any other ingredient. Further, we reveal why those hyperparameters become easier to find: as scale increases, the hyperparameter loss surface becomes lower dimensional. Nevertheless while scaling laws exist in small models, extrapolation hits statistical limitations. A holistic approach is required. Synthesizing our insights with the recent literature, we develop a new methodology for model-centric research and demonstrate it on a question that once took the field years to settle: where to place normalization layers in the transformer architecture. From small-scale experiments, we recover the large scale result: pre-normalization works better as models grow in size. With the right tools and a better understanding, small-scale experiments can deliver on scaling laws' long-awaited promise.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis