NTH

The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction

AuthorsYu Lin, Yiming Wang, Runyuan Cai, Hanze Liu, Xiaodong Zeng

September 18, 2026 2 min read
Watch on YouTube
The one-line take

Edge0 uses predicted expert routing to stream a 35B mixture-of-experts model from SSD fast enough to run on a 24GB consumer machine.

Key results

20.4
35B decode throughput

Edge0-35B tok/s on a 24 GB Apple-silicon machine

2.9
Peak active memory

GiB for the 35B tier while streaming a 19.5 GB checkpoint from SSD

82%
K=4 prerouter gain

Decode improvement from 3.5 to 6.4 tok/s on the 16 GB system

79.2
35B OpenCompass average

Served int4 model average across AIME 2026, HumanEval, GPQA-Diamond, MMLU-Pro, and IFBench

What the paper found

Edge0 targets the weight half of the memory wall: a 35B Mixture-of-Experts model such as Qwen3.6-35B-A3B still occupies 19.5 GB at int4 even though only about 3B parameters activate per token. Instead of keeping experts resident, Edge0 memory-maps them from SSD and uses a trained per-layer prerouter to predict the next layer’s expert selection one token ahead. That prediction becomes the routing itself, allowing staged SSD reads to overlap decoding without dropped experts or fallback loads. A frozen int4 model is paired with an unmerged recovery LoRA, trained on the student path with prerouter routing, supervised fine-tuning, and on-policy distillation; keeping the adapter separate avoids the degradation caused by merging and requantizing it. On a 24 GB Apple-silicon machine using MLX, the Qwen-based 35B tier reaches 20.4 tok/s with only 2.9 GiB of peak active memory, versus 3.9 tok/s and 18.2 GiB for a fully resident baseline. On a storage-stressed 16 GB system, prerouting improves K=4 decode from 3.5 to 6.4 tok/s, an 82% gain. Across AIME 2026, HumanEval, GPQA-Diamond, MMLU-Pro, and IFBench, the served 35B model averages 79.2 versus 83.2 for its fp16 teacher. The framework also supports DeepSeek-style sigmoid-group routing and an 8B Ling hybrid tier, with open checkpoints and adapters.

Original abstract

Mixture-of-experts (MoE) inference on consumer hardware is bounded by weight memory: a 35B-class model is 19.5GB at 4-bit, and sparsity shrinks the compute per token, not the bytes that must be held. Naive offloading to SSD does not help on its own, because layer N+1's experts must be chosen before layer N's output exists, so the reads cannot start early enough to hide behind compute. We present Edge0, a streaming MoE inference engine that closes the gap with a prerouter: a per-layer head predicts the next layer's routing one token ahead, and the prediction is consumed as the routing itself, so the staged expert set equals the routed set and nothing is dropped. An unmerged recovery LoRA, trained on the student path, pays back the quality lost to int4 quantization and routing replacement. On a single 24GB machine, Edge0 serves a 35B MoE at 20tok/s inside 3GiB of peak active memory, within a few points of its fp16 teacher on average across five public benchmarks. An 8B tier runs on the same framework, and the framework, checkpoints, and adapters are open source.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis