MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching
AuthorsXingwei Sun, Heinrich Dinkel, Gang Li, Jiahao Mei, Yadong Niu, Zerui Han, Yuepeng Jiang, Jiahao Zhou, Lichun Fan, Jian Luan
MiDashengLM-Gen uses an LLM and flow matching to generate coherent multilingual audio scenes with speech, music, and sound effects while sharply improving speech intelligibility.
Key results
Hours of mixed speech, music, and sound-effect data used for training.
Parameter scale of the Qwen3 backbone.
MiDashengLM-Gen’s English word error rate.
Mean error rate across nine languages with alignment pre-training.
FAD for the SM0 mixed-audio category.
Latent size that the DiT decoder width must exceed for convergence.
What the paper found
Xiaomi’s MiDashengLM-Gen is an end-to-end text-to-audio system for generating coherent scenes that combine speech, music, sound effects, and environmental sound. It replaces Dasheng AudioGen’s frozen text encoder and fixed-length diffusion pipeline with Qwen3-1.7B, which jointly models text and audio history, plus per-token conditional flow matching in the high-dimensional DashengTokenizer latent space. Audio is generated autoregressively at variable length, with a learned stop head and a 10-step Euler solver. Training uses 77k hours of mixed audio and dedicated speech data, while structured captions provide six views covering the global scene, transcript, speaker style, effects, music, and environment. On Seed-TTS, English WER falls from Dasheng AudioGen’s 12.15% to 2.79%, approaching Qwen3-TTS at 1.24%; Chinese CER is 3.87%. On multilingual evaluation, mean WER reaches 7.68%, compared with 31.73% without audio-text alignment pre-training. On MECAT, the speech-and-music mixed category achieves FAD 0.98 versus 1.70 for Dasheng AudioGen. Ablations also show that the flow-matching DiT width must exceed the 768-dimensional audio latent size for convergence, establishing an architectural scaling rule. The model remains weaker than specialized systems such as TangoFlux for isolated sound effects, but offers a unified framework for multilingual, variable-length mixed-audio generation.
Original abstract
Generating coherent audio scenes that simultaneously blend speech, music, and sound effects remains a significant challenge. Current approaches typically rely on a disjointed pipeline where a frozen, decoupled text encoder feeds a separate audio decoder, limiting cross-modal optimization and leading to poor speech intelligibility. To overcome these limitations, we introduce MiDashengLM-Gen, an end-to-end framework that couples a pre-trained Large Language Model (LLM) with per-token conditional flow matching for autoregressive, variable-length mixed-audio scene generation. MiDashengLM-Gen represents a first approach for general text-to-audio generation with one end-to-end trained model. Empirical evaluations demonstrate that MiDashengLM-Gen drastically improves speech intelligibility over existing unified models. On the Seed-TTS benchmark, English Word Error Rate (WER) drops from 12.15% to 2.79%, approaching the performance of dedicated Text-to-Speech (TTS) systems (1.24%). Furthermore, the framework extends effectively to multilingual settings, yielding highly competitive multilingual WERs compared to existing baselines. Lastly, the model maintains competitive mixed-audio generation quality on the MECAT benchmark. Code and checkpoints are available at https://github.com/xiaomi-research/midashenglm-gen and https://huggingface.co/mispeech/midashenglm-gen, and the demo page is available at https://xingws.github.io/midashenglm-gen-demo/.
Read the original paperMore in Speech AI
Browse all 27 papers →Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen
A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.
Rethinking Automated Voice Similarity by Shifting from EER to Embedding Geometry
Szu-Chi Chen, Jia-Kai Dong, Yi-Cheng Lin, Sung-Feng Huang, Hung-yi Lee
The study argues that voice-similarity systems should be judged by whether their embedding geometry matches human perception, not merely by verification accuracy.
Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
Jian Chen, You Zhang, Mark Vinton
Tacit-TTS makes zero-shot voice cloning over ten times faster while preserving the ability to clone voices from speech without transcripts, including multilingual and non-lexical references.