Unified Audio Intelligence Without Regressing on Text Intelligence
AuthorsZhifeng Kong, Sang-gil Lee, Jaehyeon Kim, Boxin Wang, Zihan Liu, Sungwon Kim, Yang Chen, Arushi Goel, Rajarshi Roy, Wenliang Dai, Zhuolin Yang, Yangyi Chen, Dongfu Jiang, Sreyan Ghosh, Tuomas Rintamaki, Andrew Tao, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, Wei Ping
Resources
This paper builds a single model that can understand and generate both audio and text, while trying hard not to lose the reasoning and alignment skills of its text-only LLM core.
Key results
total audio tokens used in training
total text tokens used in training
total input-output pairs across text and audio tasks
Audex-30B-A3B tool-integrated reasoning score
Audex-30B-A3B tool-integrated reasoning score
Audex-30B-A3B audio understanding score
What the paper found
NVIDIA’s Nemotron-Labs-Audex-30B-A3B, or Audex, is a unified audio-text large language model built on the text-only Nemotron-Cascade-2-30B-A3B backbone, designed to handle audio understanding, speech recognition, speech translation, text-to-speech, audio generation, and speech-to-speech in a single Transformer decoder without the usual text regression seen in multimodal systems. The model projects audio embeddings into the text space, then autoregressively predicts text tokens plus discrete speech and audio codec tokens, and it is trained with a large mixed corpus of 157.4B audio tokens and 320.5B text tokens across 394M samples. Audex preserves text intelligence nearly intact while improving key text-reasoning benchmarks, reaching 91.2 on AIME 2025 and 92.2 on HMMT Feb25 with tool use, while also setting strong audio results such as 75.6 on MMAU, 6.82 WER on OpenASR, 34.0 BLEU / 86.9 COMET on Fleurs speech translation, 1.70 WER on Seed-TTS-Eval fixed-voice TTS, 66.9 on AudioCaps FDopenl3, 62.7 on SongDescriber FDopenl3, and 90.0 on BigBenchAudio. The architecture uses a 30B MoE backbone with 3B activated parameters, a 52-layer hybrid Mamba-Transformer design, 128 routable experts, and a 205,312-token vocabulary after adding 65,536 speech tokens and 8,192 audio tokens. NVIDIA reports training on 512 H100 GPUs and favors a multi-stage SFT curriculum, because single-stage consolidation nearly breaks long-context behavior, collapsing NIAH from 99.3|86.8 to 6.0|0.0.
Original abstract
Audio intelligence involves understanding, reasoning about, and generating both audio and speech. In this work, we introduce Nemotron-Labs-Audex-30B-A3B (Audex), a unified audio-text LLM built on Nemotron-Cascade-2-30B-A3B, a strong text-only MoE LLM. Audex adopts a simple unified design with a single Transformer decoder: audio inputs are encoded and projected into the text embedding space, while text tokens and quantized audio output tokens are treated uniformly during generation. This architecture enables strong audio-text fusion, seamless multimodal generation, and compatibility with standard LLM training and inference infrastructure. For training, we meticulously curate audio-text datasets comprising 157.4B audio tokens and 320.5B text tokens. We apply multi-stage supervised training on these datasets, followed by text-only Cascade RL and multi-domain on-policy distillation. Audex delivers state-of-the-art audio understanding, speech recognition and translation, text-to-speech, audio generation, and speech-to-speech generation, while preserving very compelling reasoning, alignment, knowledge, long-context, and agentic capabilities of its text-only LLM backbone with marginal or no regression. We release the model checkpoints to facilitate open research.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.