NTH

Matryoshka Language Model Suites

AuthorsNathan Godey, Yoav Artzi

August 16, 2026 2 min read
Watch on YouTube
The one-line take

One nested language model can serve as several differently sized models, cutting training costs while speeding up speculative decoding.

Key results

3B
Suite composition

The main nested suite contains 500M, 1.5B, and 3B standalone sub-models.

35B
Training data

FineWeb-Edu tokens used to train the main suite.

36%
Training compute reduction

Reduction versus independently trained baselines.

53.3
3B average benchmark accuracy

Average zero-shot accuracy across seven multiple-choice benchmarks.

26%
Speculative decoding throughput gain

Maximum reported gain for the 500M-draft, 3B-verifier pair at draft length 6.

5.7%
1.5B–3B agreement improvement

Final next-token agreement improvement over independently trained models.

What the paper found

Matryoshka Language Model Suites proposes training an entire language-model family as one nested Transformer instead of running separate pretraining jobs. A 3B suite contains standalone 500M, 1.5B, and 3B sub-models, where each larger model adds width and depth to the smaller one; a norm-rescaled inter-model junction and fresh embeddings pass representations between exits. Every forward pass produces logits for all exits, enabling online knowledge distillation from the 3B teacher without separately storing teacher outputs. Trained on 35B FineWeb-Edu tokens, the suite matches independently trained Meta Llama-style baselines across ARC-Easy, ARC-Challenge, HellaSwag, LAMBADA, OpenBookQA, PIQA, and Winogrande, reaching 53.3 average accuracy at 3B with 2.067 out-of-domain byte perplexity, while reducing training compute by 36%. The nested design also shares early layers and the KV cache during speculative decoding: with a 500M draft and 3B verifier on an NVIDIA A100, throughput gains range from 14% to 26%, reaching 2,650 tokens per second at draft length 6, while cross-model next-token agreement improves by 5.7% for the 1.5B–3B pair. Compared with MatFormer, which keeps KV-cache size fixed across sub-models, Matryoshka produces detachable checkpoints whose memory footprint scales with model size, making the approach useful for deployment across devices and workloads; the authors also release a Transformers-compatible implementation for Hugging Face ecosystems.

Original abstract

Training a language model suite classically requires training each model separately and serving them independently. We improve both training and inference efficiency by stacking sub-models of increasing size into a single nested architecture trained end-to-end. This Matryoshka training framework reduces the total parameter count of the suite, enables low-cost distillation from the largest to all smaller sub-models at every training step, and is well-suited for speculative decoding as the draft model is contained within the verifier. We validate our approach by training a Matryoshka suite comprising 500M, 1.5B, and 3B sub-models. Our suite is on par with independently trained baselines on benchmark performance and validation and out-of-domain perplexities, while using 36% less training compute and improving the throughput of speculative decoding by 14-26%. We also ablate key architectural choices, offering guidance for building strong Matryoshka LM suites.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis