NTH

Luna-TTS Family Technical Report

AuthorsFeng Yin, Shuai Shi, Junjie Zheng, Kechenying Zhou, Yiqiu Wang, Chenyang He, Qiuhua Jiang, Mengxiao Bi, Yanmin Qian, Mingxin Chen, Xun Gong, Tianteng Gu, Bing Han, Peng Jiang, Chenda Li, Haiyang Sun, Han Wang, Wei Wang, Yi Wang, Leying Zhang, Wangyou Zhang, Chushu Zhou

August 18, 2026 3 min read
Watch on YouTube
The one-line take

Luna-TTS uses diffusion-style parallel speech generation to make expressive, multilingual text-to-speech faster, more controllable, and capable of voice cloning and editing.

Key results

1M
Pretraining speech corpus

Multilingual speech hours spanning Chinese, English, Japanese, and Korean.

0.6B
Backbone size

Shared Qwen3-derived parameter lineage for both Luna-TTS variants.

32
Luna-TTS refinement steps

Fixed parallel denoising steps for full-utterance generation.

0.0240
Realtime end-to-end RTF

Measured using 8-step parallel classifier-free guidance on 2 NVIDIA H20 GPUs.

41.6 ms
Realtime first-block latency

Time to produce and decode the first 1.28-second streaming audio block.

0.73
Seed-TTS-Eval Mandarin CER

Luna-TTS character error rate on the Mandarin test set.

What the paper found

Luna-TTS Family replaces conventional left-to-right codec language models with masked discrete diffusion over the full residual-vector-quantization grid. Starting from Qwen3-0.6B, the system progressively changes attention from causal to bidirectional, producing a fully non-autoregressive Luna-TTS model, then to block-causal attention for Luna-TTS Realtime. Both variants share a 0.6B-parameter lineage and are pretrained on 1M hours of Chinese, English, Japanese, and Korean speech. Luna-TTS generates complete utterances in 32 parallel refinement steps, enabling native zero-shot voice cloning and speech editing through infilling; Realtime denoises 1.28s blocks in parallel while retaining KV caching and streaming delivery. With GRPO reinforcement learning applied to realized denoising trajectories, Luna-TTS leads the compared systems on all four Seed-TTS-Eval metrics, including 0.73 CER and 79.7 SIM on Mandarin, while its expressive controls for emotion and non-verbal vocalizations compete favorably with commercial systems such as ElevenLabs and MiniMax, using Gemini 3.1 Pro Preview for model-based assessment. On NVIDIA H20 hardware, Luna-TTS Realtime reaches 0.0240 end-to-end RTF and delivers its first decoded block in 41.6 ms, illustrating a practical trade-off between global offline refinement and low-latency streaming.

Original abstract

Modern text-to-speech (TTS) is dominated by autoregressive (AR) codec language models, whose left-to-right decoding brings latency that grows with utterance length, error accumulation along the committed prefix, and an artificial generation order imposed on the Residual Vector Quantization (RVQ) token grid. We propose Luna-TTS Family, diffusion-language-model-based TTS systems pretrained on 1 million hours of speech across Chinese, English, Japanese, and Korean. The family is built by progressive adaptation of a pretrained AR text LLM, from causal to bidirectional and finally to block-causal attention, and comprises two variants sharing a single tokenizer, data pipeline, and 0.6B backbone lineage. Luna-TTS is fully non-autoregressive: it generates the entire RVQ token grid in a fixed number of parallel refinement steps, with zero-shot voice cloning and speech editing arising natively as infilling. Luna-TTS Realtime, derived by continual training, is autoregressive over blocks of 32 codec frames (1.28s) while denoising each block in parallel; it supports KV-cached blockwise generation and incremental audio delivery, achieving an end-to-end RTF of 0.0240 and 41.6 ms local first-block latency under the warmed serving protocol. An annealed fine-tuning stage adds explicit control over emotion and non-verbal vocalizations (NVVs), and a reinforcement-learning stage applies GRPO with policy ratios computed over the realized denoising trajectory. On Seed-TTS-Eval, Luna-TTS achieves the best results on all four metrics among compared open-source and commercial systems (0.73 CER / 79.7 SIM on test-zh, 1.49 WER / 76.8 SIM on test-en); on the harder in-the-wild CV3-Eval, it posts the lowest Mandarin and English error rates in our comparison. Against leading commercial systems, it achieves the best results on most objective, model-based, and human-rated metrics for NVV and emotion control.

Read the original paper

More in Speech AI

Browse all 27 papers →
01Speech

Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen

A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.

Read analysis