NTH

MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching

AuthorsXingwei Sun, Heinrich Dinkel, Gang Li, Jiahao Mei, Yadong Niu, Zerui Han, Yuepeng Jiang, Jiahao Zhou, Lichun Fan, Jian Luan

August 18, 2026 2 min read
Watch on YouTube
The one-line take

MiDashengLM-Gen uses an LLM and flow matching to generate coherent multilingual audio scenes with speech, music, and sound effects while sharply improving speech intelligibility.

Key results

77k
Training audio

Hours of mixed speech, music, and sound-effect data used for training.

1.7B
LLM backbone

Parameter scale of the Qwen3 backbone.

2.79%
Seed-TTS English WER

MiDashengLM-Gen’s English word error rate.

7.68%
Multilingual mean WER

Mean error rate across nine languages with alignment pre-training.

0.98
MECAT speech-music FAD

FAD for the SM0 mixed-audio category.

768
Audio latent dimensionality

Latent size that the DiT decoder width must exceed for convergence.

What the paper found

Xiaomi’s MiDashengLM-Gen is an end-to-end text-to-audio system for generating coherent scenes that combine speech, music, sound effects, and environmental sound. It replaces Dasheng AudioGen’s frozen text encoder and fixed-length diffusion pipeline with Qwen3-1.7B, which jointly models text and audio history, plus per-token conditional flow matching in the high-dimensional DashengTokenizer latent space. Audio is generated autoregressively at variable length, with a learned stop head and a 10-step Euler solver. Training uses 77k hours of mixed audio and dedicated speech data, while structured captions provide six views covering the global scene, transcript, speaker style, effects, music, and environment. On Seed-TTS, English WER falls from Dasheng AudioGen’s 12.15% to 2.79%, approaching Qwen3-TTS at 1.24%; Chinese CER is 3.87%. On multilingual evaluation, mean WER reaches 7.68%, compared with 31.73% without audio-text alignment pre-training. On MECAT, the speech-and-music mixed category achieves FAD 0.98 versus 1.70 for Dasheng AudioGen. Ablations also show that the flow-matching DiT width must exceed the 768-dimensional audio latent size for convergence, establishing an architectural scaling rule. The model remains weaker than specialized systems such as TangoFlux for isolated sound effects, but offers a unified framework for multilingual, variable-length mixed-audio generation.

Original abstract

Generating coherent audio scenes that simultaneously blend speech, music, and sound effects remains a significant challenge. Current approaches typically rely on a disjointed pipeline where a frozen, decoupled text encoder feeds a separate audio decoder, limiting cross-modal optimization and leading to poor speech intelligibility. To overcome these limitations, we introduce MiDashengLM-Gen, an end-to-end framework that couples a pre-trained Large Language Model (LLM) with per-token conditional flow matching for autoregressive, variable-length mixed-audio scene generation. MiDashengLM-Gen represents a first approach for general text-to-audio generation with one end-to-end trained model. Empirical evaluations demonstrate that MiDashengLM-Gen drastically improves speech intelligibility over existing unified models. On the Seed-TTS benchmark, English Word Error Rate (WER) drops from 12.15% to 2.79%, approaching the performance of dedicated Text-to-Speech (TTS) systems (1.24%). Furthermore, the framework extends effectively to multilingual settings, yielding highly competitive multilingual WERs compared to existing baselines. Lastly, the model maintains competitive mixed-audio generation quality on the MECAT benchmark. Code and checkpoints are available at https://github.com/xiaomi-research/midashenglm-gen and https://huggingface.co/mispeech/midashenglm-gen, and the demo page is available at https://xingws.github.io/midashenglm-gen-demo/.

Read the original paper

More in Speech AI

Browse all 27 papers →
01Speech

Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen

A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.

Read analysis