Small-Scale Experiments: Are We There Yet?
AuthorsNicholas Lourie, Kyunghyun Cho, Karen Ullrich, Sanae Lotfi
Resources
Small models can predict large-model behavior after all—but only when their hyperparameters are tuned carefully enough to reveal the hidden scaling laws.
Key results
Scaling laws emerge at models with 4M effective parameters when hyperparameters are thoroughly tuned.
The scaling law becomes accurate only after evaluating 256 configurations.
Learning-rate decay reduced held-out test mean-squared error by 98%.
The noisy quadratic analysis shows the effective hyperparameter count falls to 1 as models scale.
Experiments use a 100B-token subset of FineWeb-Edu.
What the paper found
This paper argues that scaling laws do exist far below production-model scale, but small models are difficult to tune, so the laws remain hidden unless researchers reach the fully optimized hyperparameter frontier. Experiments with a Llama decoder in Meta Lingua, using OpenAI’s p50k_base tokenizer, the 100B-token FineWeb-Edu subset, and NVIDIA A100 GPUs, span 4M to 268M effective parameters. Random-search studies show that reliable scaling requires far more tuning than typical practice: the relationship is absent with only 4 or 16 configurations, becomes visible at 64, and is accurate at 256 configurations. Among methodological choices, learning-rate decay reduced held-out test mean-squared error by 98%, while parameter-count conventions and tied scaling exponents mattered much less. The proposed noisy quadratic limit reveals that hyperparameter loss surfaces become lower-dimensional as models grow, with the effective hyperparameter count falling to 1, explaining why large models are easier to tune. The authors combine this diagnostic with scaling laws and perplexity-capability correspondence, validating conclusions on AI2 ARC, BoolQ, MMLU, OpenBookQA, and PIQA while warning that extrapolation eventually becomes statistically unreliable. In a transformer normalization case study, small-scale experiments recover the established result that pre-normalization scales better than post-normalization: post-norm remains more hyperparameter-sensitive, whereas pre-norm yields a cleaner scaling law and stronger compute efficiency near observed data. The practical recommendation is to explore hundreds of configurations at small scale, verify diagnostics, and transfer the resulting design upward rather than relying on naïve extrapolation.
Original abstract
Scaling laws promised cost-effective experiments; six years later, they have yet to fully deliver. Instead, researchers have found them unreliable at small scales (starting at 4M parameters) and concluded that sizable models cannot be avoided. We show this is not the case: the confounding factor is hyperparameters. Small models are highly sensitive, but hyperparameter sensitivity fades with scale. This small-scale sensitivity makes scaling laws easy to miss because they only emerge on the fully tuned frontier, and reaching that frontier requires an extensive search far beyond what most ever run. By ablating the basic scaling law recipe, we show well-tuned hyperparameters matter more than any other ingredient. Further, we reveal why those hyperparameters become easier to find: as scale increases, the hyperparameter loss surface becomes lower dimensional. Nevertheless while scaling laws exist in small models, extrapolation hits statistical limitations. A holistic approach is required. Synthesizing our insights with the recent literature, we develop a new methodology for model-centric research and demonstrate it on a question that once took the field years to settle: where to place normalization layers in the transformer architecture. From small-scale experiments, we recover the large scale result: pre-normalization works better as models grow in size. With the right tools and a better understanding, small-scale experiments can deliver on scaling laws' long-awaited promise.
Read the original paperMore in Efficient AI
Browse all 55 papers →Decoding Looped Transformers Better for (Almost) Free
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang
For LLM prefix caches, simple recency may beat fancy eviction rules, especially when workloads follow predictable session patterns.