Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs
AuthorsSong Bian, Tao Yu, Shivaram Venkataraman, Youngsuk Park
Resources
This paper shows how to use scaling laws plus architecture choices like attention design to build language models that are both cheaper to run and more accurate.
Key results
More than 200 trained models were used to fit the conditional scaling law.
Training covered models from 80M to 3B parameters.
Training data spanned 8B to 100B tokens.
Panda-1B improved average downstream accuracy over LLaMA-3.2-1B.
Panda-3B improved average downstream accuracy over LLaMA-3.2-3B.
Surefire-1B and Surefire-3B achieved up to 42% higher inference throughput on A100 with vLLM.
What the paper found
This ICLR 2026 paper from UW-Madison and Amazon Web Services extends Chinchilla-style scaling laws to include model architecture, targeting inference-efficient LLMs rather than compute-optimal training alone. The authors show that, at fixed parameter budgets, hidden size, the MLP-to-attention ratio, and grouped-query attention all materially affect throughput: larger hidden sizes, higher MLP-to-attention ratios, and higher GQA consistently increase tokens per second on both LLaMA-3.2 and Qwen3 variants. Using more than 200 decoder-only models trained from 80M to 3B parameters and 8B to 100B tokens, they fit a two-step conditional scaling law that first estimates the optimal loss under Chinchilla and then calibrates architectural effects through simple multiplicative or additive terms. The resulting predictors are accurate across scales, with low MSE and strong rank correlation, and they reveal U-shaped loss curves for both hidden size and MLP-to-attention ratio, implying interior optima rather than monotonic scaling. At 1B and 3B scales, the method yields Panda-1B and Panda-3B architectures that improve average downstream accuracy by 2.1% and 0.6% over LLaMA-3.2 baselines, while the search-derived Surefire-1B and Surefire-3B variants preserve loss constraints and raise inference throughput by up to 42% on A100 with vLLM; the same efficiency gains reach 47% with SGLang on H200. The paper also finds that fitting the law with outlier MLP-to-attention ratios below 0.5 or above 5 hurts prediction quality, and that simple separable calibration outperforms more complex non-separable formulations.
Original abstract
Scaling the number of parameters and the size of training data has proven to be an effective strategy for improving large language model (LLM) performance. Yet, as these models grow increasingly powerful and widely deployed, the cost of inference has become a pressing concern. Despite its importance, the trade-off between model accuracy and inference efficiency remains underexplored. In this work, we examine how key architectural factors, hidden size, the allocation of parameters between MLP and attention (mlp-to-attention ratio), and grouped-query attention (GQA), influence both inference cost and accuracy. We introduce a conditional scaling law that augments the Chinchilla framework with architectural information, along with a search framework for identifying architectures that are simultaneously inference-efficient and accurate. To validate our approach, we train more than 200 models spanning 80M to 3B parameters and 8B to 100B training tokens, and fit the proposed conditional scaling law. Our results show that the conditional scaling law reliably predicts optimal architectural choices and that the resulting models outperform existing open-source baselines. Under the same training budget, optimized architectures achieve up to 2.1% higher accuracy and 42% greater inference throughput compared to LLaMA-3.2.
Read the original paperMore in Efficient AI
Browse all 55 papers →Decoding Looped Transformers Better for (Almost) Free
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang
For LLM prefix caches, simple recency may beat fancy eviction rules, especially when workloads follow predictable session patterns.