NTH

AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing

AuthorsZiyang Ma, Zhikang Niu, Wenming Tu, Tianrui Wang, Ruiqi Yan, Junxi Liu, Yanru Huo, Nickk Huang, Yang Liu, Qicong Xie, Zeyu Xie, Hui Wang, Haitao Li, Zixuan Jiang, Yalin Li, Jie Fang, Yifan Duan, Zeyue Tian, Guangzheng Li, Haina Zhu, Shuyi Wang, Jinwen Wang, Mingyu Cui, Tian Tan, Auden, Sen Liang, Steve Yves, Shan Yang, Liefeng Bo, Zilong Zheng, Kai Yu, Eng-Siong Chng, Xie Chen

September 12, 2026 2 min read
Watch on YouTube
The one-line take

AuK is an open-source speech foundation model that can generate and edit audio from natural-language instructions, with a distilled version offering much faster inference.

Key results

3.03B
Training instances

Instruction–audio instances spanning five speech and audio task families.

1.95M
Effective supervision

Total effective audio-supervision hours used for unified training.

1.5B
AuK backbone size

Approximate parameters in the hybrid MMDiT–DiT Transformer backbone.

4.5
AuK-Flash speedup

Wall-clock speedup over the full model under matched inference conditions.

2.65%
Seed-TTS-Eval average error

AuK's average recognition error across the benchmark's English and Chinese subsets.

48.23%
MMAE-Speech instruction following

AuK's instruction-following rubric score on general speech editing.

What the paper found

AuK is an open-source foundational model that unifies zero-shot and instruction-controlled speech generation with content replacement, emotion and accent transformation, pitch, loudness and speed editing, enhancement, separation, and singing-lyric editing through natural-language instructions plus optional audio context. Its training corpus contains approximately 3.03 billion instruction–audio instances and 1.95 million hours of effective supervision across five task families. The architecture combines Qwen2.5-Omni semantic conditioning, Qwen3-Omni-generated annotations, a 24-kHz audio VAE, and a FLUX-style hybrid rectified-flow Transformer with dual-stream MMDiT followed by single-stream DiT blocks; the backbone has approximately 1.5 billion parameters. Human-feedback preference optimization targets open-ended editing, while Flow-GRPO reinforcement learning improves content accuracy, speaker similarity, and style consistency, using reward models based on systems such as Qwen2.5-Omni and language-model tooling that can include GPT-5.6. Distillation with consistency initialization and task-routed Decoupled DMD produces AuK-Flash, which uses 4-step, classifier-free-guidance-free inference and delivers a 4.5× wall-clock speedup over the full model. On Seed-TTS-Eval, AuK reaches 2.65% average recognition error and 0.795 speaker similarity, while on MMAE-Speech it reaches 48.23% instruction-following success and 88.11% consistency, demonstrating strong generation and general editing alongside competitive enhancement and separation.

Original abstract

We introduce AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context. To support this broad capability set, we construct approximately 3.03 billion instruction--audio instances and 1.95 million hours of effective supervision across five task families: speech generation, content editing, enhancement and separation, paralinguistic editing, and acoustic editing. AuK combines a multimodal large language model for semantic conditioning, an VAE jointly trained on speech, general audio, and music for acoustic conditioning, and a hybrid rectified-flow Transformer that performs dual-stream MMDiT blocks followed by unified single-stream DiT blocks for generation. Training begins with generation-only warm-up and proceeds to joint generation--editing pre-training. We then apply complementary post-training strategies: human-feedback preference optimization for open-ended editing and reward-based reinforcement learning for speech generation. To reduce inference cost, we further distill the model with consistency initialization and task-routed Decoupled DMD. The resulting AuK-Flash performs 4-step inference without classifier-free guidance and achieves a 4.5 wall-clock speedup over the full model under matched conditions. Experiments demonstrate leading performance on zero-shot and instruction-controlled speech generation and general instruction-guided editing, while remaining competitive on signal-level restoration tasks. We release both the source code and model weights to support reproducibility and further research.

Read the original paper

More in Speech AI

Browse all 27 papers →
01Speech

Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen

A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.

Read analysis