NTH

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs

AuthorsSong Bian, Tao Yu, Shivaram Venkataraman, Youngsuk Park

June 9, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows how to use scaling laws plus architecture choices like attention design to build language models that are both cheaper to run and more accurate.

Key results

200
model count

More than 200 trained models were used to fit the conditional scaling law.

80M-3B
parameter range

Training covered models from 80M to 3B parameters.

8B-100B
token range

Training data spanned 8B to 100B tokens.

2.1%
accuracy gain

Panda-1B improved average downstream accuracy over LLaMA-3.2-1B.

0.6%
accuracy gain

Panda-3B improved average downstream accuracy over LLaMA-3.2-3B.

42%
throughput gain

Surefire-1B and Surefire-3B achieved up to 42% higher inference throughput on A100 with vLLM.

What the paper found

This ICLR 2026 paper from UW-Madison and Amazon Web Services extends Chinchilla-style scaling laws to include model architecture, targeting inference-efficient LLMs rather than compute-optimal training alone. The authors show that, at fixed parameter budgets, hidden size, the MLP-to-attention ratio, and grouped-query attention all materially affect throughput: larger hidden sizes, higher MLP-to-attention ratios, and higher GQA consistently increase tokens per second on both LLaMA-3.2 and Qwen3 variants. Using more than 200 decoder-only models trained from 80M to 3B parameters and 8B to 100B tokens, they fit a two-step conditional scaling law that first estimates the optimal loss under Chinchilla and then calibrates architectural effects through simple multiplicative or additive terms. The resulting predictors are accurate across scales, with low MSE and strong rank correlation, and they reveal U-shaped loss curves for both hidden size and MLP-to-attention ratio, implying interior optima rather than monotonic scaling. At 1B and 3B scales, the method yields Panda-1B and Panda-3B architectures that improve average downstream accuracy by 2.1% and 0.6% over LLaMA-3.2 baselines, while the search-derived Surefire-1B and Surefire-3B variants preserve loss constraints and raise inference throughput by up to 42% on A100 with vLLM; the same efficiency gains reach 47% with SGLang on H200. The paper also finds that fitting the law with outlier MLP-to-attention ratios below 0.5 or above 5 hurts prediction quality, and that simple separable calibration outperforms more complex non-separable formulations.

Original abstract

Scaling the number of parameters and the size of training data has proven to be an effective strategy for improving large language model (LLM) performance. Yet, as these models grow increasingly powerful and widely deployed, the cost of inference has become a pressing concern. Despite its importance, the trade-off between model accuracy and inference efficiency remains underexplored. In this work, we examine how key architectural factors, hidden size, the allocation of parameters between MLP and attention (mlp-to-attention ratio), and grouped-query attention (GQA), influence both inference cost and accuracy. We introduce a conditional scaling law that augments the Chinchilla framework with architectural information, along with a search framework for identifying architectures that are simultaneously inference-efficient and accurate. To validate our approach, we train more than 200 models spanning 80M to 3B parameters and 8B to 100B training tokens, and fit the proposed conditional scaling law. Our results show that the conditional scaling law reliably predicts optimal architectural choices and that the resulting models outperform existing open-source baselines. Under the same training budget, optimized architectures achieve up to 2.1% higher accuracy and 42% greater inference throughput compared to LLaMA-3.2.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis