FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution
AuthorsShuo Yang, Xiaoze Fan, Melissa Pan, Haocheng Xi, Zhe Wang, Shanlin Sun, Kurt Keutzer, Song Han, Matei Zaharia, Chenfeng Xu, Ion Stoica
Resources
FreeToken aims to make frontier-scale mixture-of-experts models practical on ordinary personal computers by dynamically coordinating CPU, GPU, memory, and model state.
Key results
Peak decode throughput in tok/s on an RTX 5090.
Maximum decode-throughput improvement over edge baselines.
Upper bound in seconds for FreeToken across evaluated workloads.
Decode throughput in tok/s for the 35B model on an 8 GB laptop GPU.
Model size served on a single workstation GPU.
What the paper found
FreeToken is an edge-native serving system for sparse mixture-of-experts models, designed for personal machines rather than datacenter clusters. It targets models such as DeepSeek-V4-Flash, GLM-5.2, Qwen3.6-35B-A3B, and OpenAI’s GPT-OSS, while supporting real agent workloads including Claude Code and tool-using applications. Its core contribution is bandwidth-adaptive execution: a shared GPU LRU expert cache captures token-to-token routing locality, while the q⋆ policy divides cache misses between PCIe transfers and direct CPU execution according to measured host and interconnect bandwidth. During prefill, full-layer double buffering overlaps expert loading with GPU computation; semantic checkpoints at thinking and tool-call boundaries let recurrent state and prefixes survive agentic context edits, avoiding redundant recomputation. FreeToken also resizes the expert cache dynamically as VRAM and KV-cache demands change, and uses NVIDIA NVFP4 or native model quantization without modifying model computation. On an RTX 5090, it reaches up to 83 tok/s on Qwen3.6-35B-A3B and achieves as much as 2.3× the decode throughput of edge baselines including llama.cpp, Ollama, and KTransformers. Its worst-case time to first token remains below 44 s, compared with baseline stalls above 150 s. The system serves a 35B model at 39.3 tok/s on an 8 GB RTX 4060 laptop, and runs the 753B GLM-5.2 at workstation scale, demonstrating that bandwidth-aware software can bring frontier open-weight models to consumer hardware.
Original abstract
Frontier open-weight models are increasingly available, but serving them still largely assumes datacenter infrastructure. We present FreeToken, an edge-native MoE serving system that treats a personal machine not as a small GPU, but as a unified, elastic inference platform. FreeToken co-designs the full serving stack, including model layout and loading, expert residency, CPU--GPU execution, agentic state reuse, and runtime memory management, around two realities of local AI: agent workloads continuously change their execution pattern, and edge hardware exposes heterogeneous resources whose balance differs from machine to machine. Rather than committing to a fixed offloading strategy, FreeToken continuously maps computation and model state onto the resources actually available. FreeToken supports more than 20 MoE models and real coding and tool-using agents across hardware ranging from an 8GB laptop GPU to a single workstation GPU. More importantly, it changes what these machines can practically serve, from a 35B model on a laptop to a 284B model on a gaming desktop and the 753B GLM-5.2 on a single workstation GPU. FreeToken turns open weights into deployable local software, making the machines users already own a practical platform for frontier-scale intelligence. We release the system at flashml.ai.
Read the original paperMore in Efficient AI
Browse all 55 papers →Decoding Looped Transformers Better for (Almost) Free
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang
For LLM prefix caches, simple recency may beat fancy eviction rules, especially when workloads follow predictable session patterns.