Qwen-Music Technical Report
AuthorsJin Xu, Kangdi Wang, Ruibin Yuan, Shun Lei, Xiong Wang, Xize Cheng, Xueyao Zhang, Yang Zhang, Yiheng Chen, Yongqi Wang, Yue Wang, Zhifang Guo, Zihan Liu, Zijian Lin, Dake Guo, Hangrui Hu, Lei Xie, Linhan Ma, Wei Xue, Wenxiang Guo, Xinfa Zhu, Xipin Wei, Yangze Li, Yuanjun Lv, Yuxuan Wang, Yunfei Chu, Zhiyong Wu
Resources
Qwen-Music combines melody planning, large-scale multilingual training, and waveform rendering to generate and reinterpret high-quality songs with vocals.
Key results
Qwen-Music-Tokenizer produces Music Semantic Tokens at 25 Hz.
The LLM is trained on more than 5M hours of multilingual music.
Qwen-Music leads 13 of 16 objective musicality and audio-quality metrics.
Professional raters prefer Qwen-Music in 50.3% of blind comparisons with Suno V5.5.
Section-level Melody-CoT achieves a Melody MAE of 1.48 on the AI-generated reference set.
Spec-VAE with the Band-Mode Refiner achieves 0.461 Mel Distance on the Song Describer Dataset.
What the paper found
The Qwen Team introduces Qwen-Music, a unified system for text-to-song and reference-based cover generation, built from a 25 Hz single-codebook semantic tokenizer, an autoregressive Qwen-Music-LLM initialized from the 3B Qwen3.5-Omni, and a generative renderer producing 48 kHz stereo audio. Its central innovation is Melody-CoT, which plans coarse vocal melodies before generating full musical-token sequences, enabling stronger long-range structure and melody cloning while still allowing changes to genre, arrangement, vocal timbre, and gender. Qwen-Music-Render combines a 1.3B Diffusion Transformer, Spec-VAE, and a frequency-aware Band-Mode Refiner to restore phase, magnitude, and stereo detail lost during semantic compression. Training uses more than 5M hours of multilingual music, a quality-graded curriculum, supervised fine-tuning, offline DPO, and online GSPO. On 600 Chinese and English prompts, Qwen-Music leads 13 of 16 objective musicality and audio-quality metrics; professional raters give it a 50.3% win rate against Suno V5.5, alongside higher win rates against Suno V5, Mureka V8, and MiniMax Music. For cover generation, it reaches a Melody MAE of 1.48 on the AI-generated reference set, while the renderer’s Band-Mode Refiner reduces Mel Distance to 0.461 on the 546-track Song Describer Dataset, demonstrating improvements in both melody preservation and acoustic reconstruction.
Original abstract
In this report, we introduce Qwen-Music, a powerful music generation model capable of producing highly musical and high-fidelity songs with complete vocal singing. Qwen-Music supports two core tasks: Text to Music Generation, which create entirely new songs from text descriptions, lyrics, and musical attributes, and Cover Song Generation, which reinterprets existing songs with different styles and vocal characteristics. Architecturally, Qwen-Music integrates three core components: Qwen-Music-Tokenizer, Qwen-Music-LLM, and Qwen-Music-Render. Qwen-Music-Tokenizer compresses audio into a 25 Hz single-codebook stream of Music Semantic Tokens that preserve semantic and melodic information for LLM prediction. Based on these tokens, Qwen-Music-LLM performs autoregressive music semantic modeling, with a key novelty being a melody-token-based chain-of-thought (Melody-CoT) mechanism that plans melodies before full-song generation, improving creativity, musicality, structural coherence, and reference-audio-based melody cloning. To overcome the fidelity limitations of discrete semantic tokens, Qwen-Music-Render performs generative stereo rendering, enriching acoustic details and producing high-fidelity stereo waveforms. Finally, we train Qwen-Music-LLM on more than 5 million hours of multilingual music data covering hundreds of languages. We first apply quality-aware pre-training curriculum, then use progressive post-training, comprising supervised initialization, offline DPO, and online GSPO, to further improve musicality and instruction-following ability. Across 600 Chinese and English prompts, Qwen-Music achieves state-of-the-art results in 13 of 16 objective musicality and audio-quality metrics. Professional evaluators also prefer Qwen-Music over leading proprietary systems. For cover song generation, Qwen-Music preserves reference melodies more accurately than leading proprietary systems.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.