AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing
AuthorsZiyang Ma, Zhikang Niu, Wenming Tu, Tianrui Wang, Ruiqi Yan, Junxi Liu, Yanru Huo, Nickk Huang, Yang Liu, Qicong Xie, Zeyu Xie, Hui Wang, Haitao Li, Zixuan Jiang, Yalin Li, Jie Fang, Yifan Duan, Zeyue Tian, Guangzheng Li, Haina Zhu, Shuyi Wang, Jinwen Wang, Mingyu Cui, Tian Tan, Auden, Sen Liang, Steve Yves, Shan Yang, Liefeng Bo, Zilong Zheng, Kai Yu, Eng-Siong Chng, Xie Chen
Resources
AuK is an open-source speech foundation model that can generate and edit audio from natural-language instructions, with a distilled version offering much faster inference.
Key results
Instruction–audio instances spanning five speech and audio task families.
Total effective audio-supervision hours used for unified training.
Approximate parameters in the hybrid MMDiT–DiT Transformer backbone.
Wall-clock speedup over the full model under matched inference conditions.
AuK's average recognition error across the benchmark's English and Chinese subsets.
AuK's instruction-following rubric score on general speech editing.
What the paper found
AuK is an open-source foundational model that unifies zero-shot and instruction-controlled speech generation with content replacement, emotion and accent transformation, pitch, loudness and speed editing, enhancement, separation, and singing-lyric editing through natural-language instructions plus optional audio context. Its training corpus contains approximately 3.03 billion instruction–audio instances and 1.95 million hours of effective supervision across five task families. The architecture combines Qwen2.5-Omni semantic conditioning, Qwen3-Omni-generated annotations, a 24-kHz audio VAE, and a FLUX-style hybrid rectified-flow Transformer with dual-stream MMDiT followed by single-stream DiT blocks; the backbone has approximately 1.5 billion parameters. Human-feedback preference optimization targets open-ended editing, while Flow-GRPO reinforcement learning improves content accuracy, speaker similarity, and style consistency, using reward models based on systems such as Qwen2.5-Omni and language-model tooling that can include GPT-5.6. Distillation with consistency initialization and task-routed Decoupled DMD produces AuK-Flash, which uses 4-step, classifier-free-guidance-free inference and delivers a 4.5× wall-clock speedup over the full model. On Seed-TTS-Eval, AuK reaches 2.65% average recognition error and 0.795 speaker similarity, while on MMAE-Speech it reaches 48.23% instruction-following success and 88.11% consistency, demonstrating strong generation and general editing alongside competitive enhancement and separation.
Original abstract
We introduce AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context. To support this broad capability set, we construct approximately 3.03 billion instruction--audio instances and 1.95 million hours of effective supervision across five task families: speech generation, content editing, enhancement and separation, paralinguistic editing, and acoustic editing. AuK combines a multimodal large language model for semantic conditioning, an VAE jointly trained on speech, general audio, and music for acoustic conditioning, and a hybrid rectified-flow Transformer that performs dual-stream MMDiT blocks followed by unified single-stream DiT blocks for generation. Training begins with generation-only warm-up and proceeds to joint generation--editing pre-training. We then apply complementary post-training strategies: human-feedback preference optimization for open-ended editing and reward-based reinforcement learning for speech generation. To reduce inference cost, we further distill the model with consistency initialization and task-routed Decoupled DMD. The resulting AuK-Flash performs 4-step inference without classifier-free guidance and achieves a 4.5 wall-clock speedup over the full model under matched conditions. Experiments demonstrate leading performance on zero-shot and instruction-controlled speech generation and general instruction-guided editing, while remaining competitive on signal-level restoration tasks. We release both the source code and model weights to support reproducibility and further research.
Read the original paperMore in Speech AI
Browse all 27 papers →Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen
A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.
Rethinking Automated Voice Similarity by Shifting from EER to Embedding Geometry
Szu-Chi Chen, Jia-Kai Dong, Yi-Cheng Lin, Sung-Feng Huang, Hung-yi Lee
The study argues that voice-similarity systems should be judged by whether their embedding geometry matches human perception, not merely by verification accuracy.
Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
Jian Chen, You Zhang, Mark Vinton
Tacit-TTS makes zero-shot voice cloning over ten times faster while preserving the ability to clone voices from speech without transcripts, including multilingual and non-lexical references.