Matryoshka Language Model Suites
AuthorsNathan Godey, Yoav Artzi
Resources
One nested language model can serve as several differently sized models, cutting training costs while speeding up speculative decoding.
Key results
The main nested suite contains 500M, 1.5B, and 3B standalone sub-models.
FineWeb-Edu tokens used to train the main suite.
Reduction versus independently trained baselines.
Average zero-shot accuracy across seven multiple-choice benchmarks.
Maximum reported gain for the 500M-draft, 3B-verifier pair at draft length 6.
Final next-token agreement improvement over independently trained models.
What the paper found
Matryoshka Language Model Suites proposes training an entire language-model family as one nested Transformer instead of running separate pretraining jobs. A 3B suite contains standalone 500M, 1.5B, and 3B sub-models, where each larger model adds width and depth to the smaller one; a norm-rescaled inter-model junction and fresh embeddings pass representations between exits. Every forward pass produces logits for all exits, enabling online knowledge distillation from the 3B teacher without separately storing teacher outputs. Trained on 35B FineWeb-Edu tokens, the suite matches independently trained Meta Llama-style baselines across ARC-Easy, ARC-Challenge, HellaSwag, LAMBADA, OpenBookQA, PIQA, and Winogrande, reaching 53.3 average accuracy at 3B with 2.067 out-of-domain byte perplexity, while reducing training compute by 36%. The nested design also shares early layers and the KV cache during speculative decoding: with a 500M draft and 3B verifier on an NVIDIA A100, throughput gains range from 14% to 26%, reaching 2,650 tokens per second at draft length 6, while cross-model next-token agreement improves by 5.7% for the 1.5B–3B pair. Compared with MatFormer, which keeps KV-cache size fixed across sub-models, Matryoshka produces detachable checkpoints whose memory footprint scales with model size, making the approach useful for deployment across devices and workloads; the authors also release a Transformers-compatible implementation for Hugging Face ecosystems.
Original abstract
Training a language model suite classically requires training each model separately and serving them independently. We improve both training and inference efficiency by stacking sub-models of increasing size into a single nested architecture trained end-to-end. This Matryoshka training framework reduces the total parameter count of the suite, enables low-cost distillation from the largest to all smaller sub-models at every training step, and is well-suited for speculative decoding as the draft model is contained within the verifier. We validate our approach by training a Matryoshka suite comprising 500M, 1.5B, and 3B sub-models. Our suite is on par with independently trained baselines on benchmark performance and validation and out-of-domain perplexities, while using 36% less training compute and improving the throughput of speculative decoding by 14-26%. We also ablate key architectural choices, offering guidance for building strong Matryoshka LM suites.
Read the original paperMore in Efficient AI
Browse all 55 papers →Decoding Looped Transformers Better for (Almost) Free
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang
For LLM prefix caches, simple recency may beat fancy eviction rules, especially when workloads follow predictable session patterns.