NTH

FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution

AuthorsShuo Yang, Xiaoze Fan, Melissa Pan, Haocheng Xi, Zhe Wang, Shanlin Sun, Kurt Keutzer, Song Han, Matei Zaharia, Chenfeng Xu, Ion Stoica

August 25, 2026 2 min read
Watch on YouTube
The one-line take

FreeToken aims to make frontier-scale mixture-of-experts models practical on ordinary personal computers by dynamically coordinating CPU, GPU, memory, and model state.

Key results

83
Qwen3.6 peak decode

Peak decode throughput in tok/s on an RTX 5090.

2.3
Baseline throughput gain

Maximum decode-throughput improvement over edge baselines.

44
Worst-case TTFT

Upper bound in seconds for FreeToken across evaluated workloads.

39.3
RTX 4060 laptop decode

Decode throughput in tok/s for the 35B model on an 8 GB laptop GPU.

753B
GLM-5.2 model size

Model size served on a single workstation GPU.

What the paper found

FreeToken is an edge-native serving system for sparse mixture-of-experts models, designed for personal machines rather than datacenter clusters. It targets models such as DeepSeek-V4-Flash, GLM-5.2, Qwen3.6-35B-A3B, and OpenAI’s GPT-OSS, while supporting real agent workloads including Claude Code and tool-using applications. Its core contribution is bandwidth-adaptive execution: a shared GPU LRU expert cache captures token-to-token routing locality, while the q⋆ policy divides cache misses between PCIe transfers and direct CPU execution according to measured host and interconnect bandwidth. During prefill, full-layer double buffering overlaps expert loading with GPU computation; semantic checkpoints at thinking and tool-call boundaries let recurrent state and prefixes survive agentic context edits, avoiding redundant recomputation. FreeToken also resizes the expert cache dynamically as VRAM and KV-cache demands change, and uses NVIDIA NVFP4 or native model quantization without modifying model computation. On an RTX 5090, it reaches up to 83 tok/s on Qwen3.6-35B-A3B and achieves as much as 2.3× the decode throughput of edge baselines including llama.cpp, Ollama, and KTransformers. Its worst-case time to first token remains below 44 s, compared with baseline stalls above 150 s. The system serves a 35B model at 39.3 tok/s on an 8 GB RTX 4060 laptop, and runs the 753B GLM-5.2 at workstation scale, demonstrating that bandwidth-aware software can bring frontier open-weight models to consumer hardware.

Original abstract

Frontier open-weight models are increasingly available, but serving them still largely assumes datacenter infrastructure. We present FreeToken, an edge-native MoE serving system that treats a personal machine not as a small GPU, but as a unified, elastic inference platform. FreeToken co-designs the full serving stack, including model layout and loading, expert residency, CPU--GPU execution, agentic state reuse, and runtime memory management, around two realities of local AI: agent workloads continuously change their execution pattern, and edge hardware exposes heterogeneous resources whose balance differs from machine to machine. Rather than committing to a fixed offloading strategy, FreeToken continuously maps computation and model state onto the resources actually available. FreeToken supports more than 20 MoE models and real coding and tool-using agents across hardware ranging from an 8GB laptop GPU to a single workstation GPU. More importantly, it changes what these machines can practically serve, from a 35B model on a laptop to a 284B model on a gaming desktop and the 753B GLM-5.2 on a single workstation GPU. FreeToken turns open weights into deployable local software, making the machines users already own a practical platform for frontier-scale intelligence. We release the system at flashml.ai.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis