NTH

On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability

AuthorsZihan Qiu, Zekun Wang, Xiao Li, Yanpeng Li, Yang Xu, Yixuan Wang, Huaqing Zhang, Rui Men, Bochao Mao, Chengruidong Zhang, Fan Zhou, Hao Luo, Haofeng Huang, Haoran Lian, Haoyan Huang, Hongqing Chen, Jianwei Zhang, Jing Xu, Junjie Wang, Langshi Chen, Liangyu Wang, Linlang Jiang, Man Yuan, Minmin Sun, Peng Jin, Siqi Zhang, Siyu Wang, Xingzhang Ren, Yakai Wang, Yi Zhang, Yiming Dong, Yizhong Cao, Yubo Ma, Yunfei Mao, Bo Zheng, Dayiheng Liu

September 3, 2026 3 min read
Watch on YouTube
The one-line take

Qwen3.8-Next combines sparse experts, hybrid attention, external memory, and optimizer changes to achieve competitive performance with far less computation and improved training stability.

Key results

125B
Total parameters

Total parameters in Qwen3.8-Flash-Next

6B
Activated parameters

Parameters activated per token

51B
N-gram embedding parameters

Parameters stored off accelerator in host-memory embedding tables

7.6
QSA prefill speedup

Kernel-level speedup over dense attention at a 1M-token context

4.9
QSA decode speedup

Kernel-level speedup over dense attention at a 1M-token context

18.8%
Batch-size warmup overhead

Additional optimizer steps required by batch-size warmup

What the paper found

Qwen3.8-Flash-Next is a 125B-parameter sparse mixture-of-experts model that activates 6B parameters per token and adds 51B host-memory n-gram embedding parameters. Its hybrid token mixer combines three Gated DeltaNet layers with one global-attention layer, while Qwen Sparse Attention, trained through dense distillation and sparse adaptation, replaces full attention during continued pretraining. At a 1M-token context, QSA delivers 7.6× faster prefill and 4.9× faster decode than dense attention, while raising the short-context benchmark average from 75.9 to 76.8 and improving long-context retrieval on OpenAI’s MRCR, including 40.53 versus 30.66 at 512K. The four-branch Gated Residual mechanism uses elementwise sigmoid gating to improve capacity, reduce memory traffic through FP8 residual storage, and stabilize optimization alongside Muon; at four times the optimal learning rate, the new recipe remains stable while the earlier Qwen3.5 structure produces 183 spikes per 10K steps. A revised scaling law favors larger batches and learning rates: batch-size warmup is discarded because it costs 18.8% more optimizer steps without improving final loss. Overall, Qwen3.8-Flash-Next matches or exceeds the much larger Qwen3.7-Plus on 8 of 14 benchmarks while using roughly one-third the activated parameters and training tokens and about one-ninth the training FLOPs, extending efficiency ideas also explored in Google DeepMind’s Gemma 3n.

Original abstract

We describe the architecture and ablations of Qwen3.8-Flash-Next, a sparse mixture-of-experts model with 125B parameters, 6B activated per token, and additional 51B parameters of n-gram embedding tables held off the accelerator. On fourteen pre-training benchmarks the model leads the 397B-A17B predecessor on eight and trails it on the rest by at most 2.6 points, at 1/3 the activated parameters, 1/3 the training tokens, and roughly 1/9 the training FLOPs. Token mixing uses a layer-wise hybrid of Gated DeltaNet (GDN) and global attention, with one full-attention layer in every four; at continued-pretraining time those full-attention layers are replaced by Qwen Sparse Attention (QSA), which scores context at micro-block granularity with a compressed lightweight indexer. The residual stream is widened to four branches and read through an elementwise gate, a design we call the Gated Residual (GR). Capacity is added outside the backbone by a single n-gram embedding layer whose tables are prefetched from host memory. We evaluate every candidate change along three axes: loss together with downstream benchmarks; the cost of the change in training, prefill and decode; and its effect on the optimal hyperparameters and training stability. Loss and downstream accuracy do not always move together: enlarging the n-gram vocabulary lowers loss monotonically while downstream accuracy saturates. The architecture and the Muon optimizer together shift the optimal learning rate and batch size upwards, render batch-size warmup unnecessary, and substantially improve stability under stress tests. Loss, benchmarks, efficiency and stability form one design problem. Solved jointly, they yield a recipe that is simultaneously more efficient, more capable and more stable.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis