YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
AuthorsRuibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
Resources
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.
Key results
Approximate parameter count of the unified AR–NAR Mixture-of-Transformers.
Hours of music used to train the joint symbolic-and-audio generator.
YuE2 score under the standard candidate-selection setting.
SongBench Global Avg achieved by selecting from eight YuE2 candidates.
Expert overall-quality preference for symbolic planning, versus 34.6% without planning.
MARBLE metrics where MERT2 surpasses previous best results, out of 15.
What the paper found
YuE2 unifies symbolic composition and full-song audio generation in one AR–NAR Mixture-of-Transformers. Given text and lyrics, its 28-layer, approximately 3.58B-parameter model first generates an editable ABC score containing melody, chords, key, meter, tempo, and form; it then predicts 25-Hz semantic music tokens and continuous acoustic latents before decoding 48-kHz stereo audio through a variational autoencoder. Trained on 346,000 hours of music, YuE2 uses MERT2 for semantic supervision and SheetSage2 for automatically transcribed lead sheets, avoiding the need for aligned scores. On the 192-prompt WildSongBench, it reaches 6.7316 SongBench Global Avg, while selecting the best of eight candidates raises the score to 6.9632, exceeding evaluated public systems. Expert listeners favor symbolic planning over direct audio generation for overall quality, 49.3% versus 34.6%, and prefer the unified model over a separate language model plus diffusion Transformer, including the Qwen-Music-style LM+DiT baseline. The same checkpoint supports score editing and zero-shot covers: edits achieve 84.17% exact target-pitch accuracy and 79.54% chord agreement while preserving most unchanged content. MERT2 leads 14 of 15 MARBLE metrics, and SheetSage2-AR leads 12 of 15 transcription benchmark–metric pairs. In comparisons with proprietary systems, YuE2’s best-of-8 setting is preferred over Suno v4.5 and is nearly balanced against Suno v5, while its average tie-adjusted audio-quality preference across six proprietary generators is 58.9%.
Original abstract
Symbolic models make melody, harmony, rhythm, and form explicit but typically stop before a finished recording; audio models produce complete songs while leaving composition implicit. We introduce YuE2, which unifies symbolic and audio music generation at frontier quality through symbolic planning. A single AR-NAR Mixture-of-Transformers (MoT) first writes a readable score specifying melody and harmony, expands it into semantic music tokens, and realizes it as full-song audio. In comparisons using the same checkpoint, experts prefer symbolic planning for overall quality and musicality, with 49.3% of overall preferences versus 34.6% without planning. Experts also favor the unified model over a separate language model and diffusion Transformer. On WildSongBench, YuE2 scores 6.73 on SongBench Global Avg, exceeding all evaluated public baselines. Selecting from eight candidates (best-of-8), YuE2 reaches 6.96, the highest observed mean among all evaluated systems. Expert listening further establishes its competitiveness with proprietary song generators, favoring best-of-8 over Suno v4.5 and yielding nearly balanced preferences against Suno v5. To learn this generation process from recordings without aligned scores, we introduce MERT2 and SheetSage2 to supply semantic and symbolic supervision. MERT2 sets a new state of the art in music representation learning, surpassing previous best results on 14 of 15 MARBLE metrics; SheetSage2 leads 12 of 15 benchmark-metric pairs in our lead-sheet transcription comparison. The same checkpoint follows score edits while largely preserving unedited musical content and generates zero-shot covers without cover-specific training. Its readable score also enables agentic music editing, with external language models translating user feedback into revisions of the composition.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs
Sebastián Andrés Cajas Ordóñez, Maximin Lange, Quang Bui, Anqi Peter Li, Felipe Ocampo Osorio, Rafi Al Attrach, Kushul Reddy Palakala, Sahil Kapadia, Zakaria Laouabdia Sellami, Xinyue Zhang, Ashley Zhang, Leo Anthony Celi
ModaLens tests whether medical VLMs truly look at the X-ray when they already have the report, revealing that reports can substantially suppress measurable image sensitivity.