NTH

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

AuthorsRuibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

October 2, 2026 3 min read
Watch on YouTube
The one-line take

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Key results

3.58B
YuE2 model size

Approximate parameter count of the unified AR–NAR Mixture-of-Transformers.

346,000
YuE2 training data

Hours of music used to train the joint symbolic-and-audio generator.

6.7316
WildSongBench Global Avg

YuE2 score under the standard candidate-selection setting.

6.9632
WildSongBench best-of-8

SongBench Global Avg achieved by selecting from eight YuE2 candidates.

49.3%
Planning preference

Expert overall-quality preference for symbolic planning, versus 34.6% without planning.

14
MERT2 MARBLE leadership

MARBLE metrics where MERT2 surpasses previous best results, out of 15.

What the paper found

YuE2 unifies symbolic composition and full-song audio generation in one AR–NAR Mixture-of-Transformers. Given text and lyrics, its 28-layer, approximately 3.58B-parameter model first generates an editable ABC score containing melody, chords, key, meter, tempo, and form; it then predicts 25-Hz semantic music tokens and continuous acoustic latents before decoding 48-kHz stereo audio through a variational autoencoder. Trained on 346,000 hours of music, YuE2 uses MERT2 for semantic supervision and SheetSage2 for automatically transcribed lead sheets, avoiding the need for aligned scores. On the 192-prompt WildSongBench, it reaches 6.7316 SongBench Global Avg, while selecting the best of eight candidates raises the score to 6.9632, exceeding evaluated public systems. Expert listeners favor symbolic planning over direct audio generation for overall quality, 49.3% versus 34.6%, and prefer the unified model over a separate language model plus diffusion Transformer, including the Qwen-Music-style LM+DiT baseline. The same checkpoint supports score editing and zero-shot covers: edits achieve 84.17% exact target-pitch accuracy and 79.54% chord agreement while preserving most unchanged content. MERT2 leads 14 of 15 MARBLE metrics, and SheetSage2-AR leads 12 of 15 transcription benchmark–metric pairs. In comparisons with proprietary systems, YuE2’s best-of-8 setting is preferred over Suno v4.5 and is nearly balanced against Suno v5, while its average tie-adjusted audio-quality preference across six proprietary generators is 58.9%.

Original abstract

Symbolic models make melody, harmony, rhythm, and form explicit but typically stop before a finished recording; audio models produce complete songs while leaving composition implicit. We introduce YuE2, which unifies symbolic and audio music generation at frontier quality through symbolic planning. A single AR-NAR Mixture-of-Transformers (MoT) first writes a readable score specifying melody and harmony, expands it into semantic music tokens, and realizes it as full-song audio. In comparisons using the same checkpoint, experts prefer symbolic planning for overall quality and musicality, with 49.3% of overall preferences versus 34.6% without planning. Experts also favor the unified model over a separate language model and diffusion Transformer. On WildSongBench, YuE2 scores 6.73 on SongBench Global Avg, exceeding all evaluated public baselines. Selecting from eight candidates (best-of-8), YuE2 reaches 6.96, the highest observed mean among all evaluated systems. Expert listening further establishes its competitiveness with proprietary song generators, favoring best-of-8 over Suno v4.5 and yielding nearly balanced preferences against Suno v5. To learn this generation process from recordings without aligned scores, we introduce MERT2 and SheetSage2 to supply semantic and symbolic supervision. MERT2 sets a new state of the art in music representation learning, surpassing previous best results on 14 of 15 MARBLE metrics; SheetSage2 leads 12 of 15 benchmark-metric pairs in our lead-sheet transcription comparison. The same checkpoint follows score edits while largely preserving unedited musical content and generates zero-shot covers without cover-specific training. Its readable score also enables agentic music editing, with external language models translating user feedback into revisions of the composition.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs

Sebastián Andrés Cajas Ordóñez, Maximin Lange, Quang Bui, Anqi Peter Li, Felipe Ocampo Osorio, Rafi Al Attrach, Kushul Reddy Palakala, Sahil Kapadia, Zakaria Laouabdia Sellami, Xinyue Zhang, Ashley Zhang, Leo Anthony Celi

ModaLens tests whether medical VLMs truly look at the X-ray when they already have the report, revealing that reports can substantially suppress measurable image sensitivity.

Read analysis