NTH

Qwen-Music Technical Report

AuthorsJin Xu, Kangdi Wang, Ruibin Yuan, Shun Lei, Xiong Wang, Xize Cheng, Xueyao Zhang, Yang Zhang, Yiheng Chen, Yongqi Wang, Yue Wang, Zhifang Guo, Zihan Liu, Zijian Lin, Dake Guo, Hangrui Hu, Lei Xie, Linhan Ma, Wei Xue, Wenxiang Guo, Xinfa Zhu, Xipin Wei, Yangze Li, Yuanjun Lv, Yuxuan Wang, Yunfei Chu, Zhiyong Wu

July 25, 2026 2 min read
Watch on YouTube
The one-line take

Qwen-Music combines melody planning, large-scale multilingual training, and waveform rendering to generate and reinterpret high-quality songs with vocals.

Key results

25
Semantic token frame rate

Qwen-Music-Tokenizer produces Music Semantic Tokens at 25 Hz.

5M
Training corpus

The LLM is trained on more than 5M hours of multilingual music.

13
Best objective metrics

Qwen-Music leads 13 of 16 objective musicality and audio-quality metrics.

50.3%
Win rate versus Suno V5.5

Professional raters prefer Qwen-Music in 50.3% of blind comparisons with Suno V5.5.

1.48
Cover melody MAE

Section-level Melody-CoT achieves a Melody MAE of 1.48 on the AI-generated reference set.

0.461
Refined Mel Distance

Spec-VAE with the Band-Mode Refiner achieves 0.461 Mel Distance on the Song Describer Dataset.

What the paper found

The Qwen Team introduces Qwen-Music, a unified system for text-to-song and reference-based cover generation, built from a 25 Hz single-codebook semantic tokenizer, an autoregressive Qwen-Music-LLM initialized from the 3B Qwen3.5-Omni, and a generative renderer producing 48 kHz stereo audio. Its central innovation is Melody-CoT, which plans coarse vocal melodies before generating full musical-token sequences, enabling stronger long-range structure and melody cloning while still allowing changes to genre, arrangement, vocal timbre, and gender. Qwen-Music-Render combines a 1.3B Diffusion Transformer, Spec-VAE, and a frequency-aware Band-Mode Refiner to restore phase, magnitude, and stereo detail lost during semantic compression. Training uses more than 5M hours of multilingual music, a quality-graded curriculum, supervised fine-tuning, offline DPO, and online GSPO. On 600 Chinese and English prompts, Qwen-Music leads 13 of 16 objective musicality and audio-quality metrics; professional raters give it a 50.3% win rate against Suno V5.5, alongside higher win rates against Suno V5, Mureka V8, and MiniMax Music. For cover generation, it reaches a Melody MAE of 1.48 on the AI-generated reference set, while the renderer’s Band-Mode Refiner reduces Mel Distance to 0.461 on the 546-track Song Describer Dataset, demonstrating improvements in both melody preservation and acoustic reconstruction.

Original abstract

In this report, we introduce Qwen-Music, a powerful music generation model capable of producing highly musical and high-fidelity songs with complete vocal singing. Qwen-Music supports two core tasks: Text to Music Generation, which create entirely new songs from text descriptions, lyrics, and musical attributes, and Cover Song Generation, which reinterprets existing songs with different styles and vocal characteristics. Architecturally, Qwen-Music integrates three core components: Qwen-Music-Tokenizer, Qwen-Music-LLM, and Qwen-Music-Render. Qwen-Music-Tokenizer compresses audio into a 25 Hz single-codebook stream of Music Semantic Tokens that preserve semantic and melodic information for LLM prediction. Based on these tokens, Qwen-Music-LLM performs autoregressive music semantic modeling, with a key novelty being a melody-token-based chain-of-thought (Melody-CoT) mechanism that plans melodies before full-song generation, improving creativity, musicality, structural coherence, and reference-audio-based melody cloning. To overcome the fidelity limitations of discrete semantic tokens, Qwen-Music-Render performs generative stereo rendering, enriching acoustic details and producing high-fidelity stereo waveforms. Finally, we train Qwen-Music-LLM on more than 5 million hours of multilingual music data covering hundreds of languages. We first apply quality-aware pre-training curriculum, then use progressive post-training, comprising supervised initialization, offline DPO, and online GSPO, to further improve musicality and instruction-following ability. Across 600 Chinese and English prompts, Qwen-Music achieves state-of-the-art results in 13 of 16 objective musicality and audio-quality metrics. Professional evaluators also prefer Qwen-Music over leading proprietary systems. For cover song generation, Qwen-Music preserves reference melodies more accurately than leading proprietary systems.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis