NTH

Multimodal Music Recommendation System using LLMs

AuthorsSrikar Prabhas Kandagatla, Sreehitha R. Narayana, Chandana Magapu, Swetha Mohan, Shamanth Kuthpadi, Hongjie Chen, Ryan A. Rossi, Franck Dernoncourt, Nesreen Ahmed

June 30, 2026 2 min read
Watch on YouTube
The one-line take

This paper builds a music recommender that uses song audio, lyrics, and AI-generated descriptions together with LLMs to better predict what people will want to hear next.

Key results

814
users

processed LastFM-1K cohort after filtering

4.21M
listening events

total scrobbles in the benchmark

421,396
sessions

sessionized interaction sequences

50,029
catalog size

top-50k song catalog size

95%
Recall improvement

best reported gain over ID-only baselines

79%
NDCG improvement

best reported gain over ID-only baselines

What the paper found

This paper, developed by researchers from the University of Massachusetts Amherst with industry collaborators from Dolby Laboratories, Adobe Research, and Cisco Research, introduces a multimodal session-based music recommender that extends the E4SRec framework with audio, lyrics, LLM-generated musicological metadata, and listening completion ratios on a new LastFM-1K benchmark. The authors curate 4.21M listening events from 814 users into 421,396 sessions and a 50,029-track catalog, then enrich each track with EnCodec, CLAP, MERT, Music2Vec, and MFCC audio features, MPNet and other lyric encoders, and 58 MGPHot perceptual attributes generated with Azure OpenAI GPT-5. Their validation shows the LLM annotations preserve ground-truth ranking structure, with category-level Spearman correlations around 0.56 for Lyrics and 0.53 for Sonority, but lower agreement for Rhythm and Composition near 0.30. Across LLaMA-3-70B, LLaMA-2-13B, and Qwen2.5-7B-Instruct, content-aware modeling outperforms ID-only baselines by up to 95% in Recall and 79% in NDCG, although the gains are highly sensitive to backbone and fusion choice. The strongest fine-tuned result reaches Recall@20 of 0.404 and NDCG@20 of 0.177 with LLaMA-2-13B plus SASRec, while zero-shot performance is more brittle and sometimes degrades under naive multimodal fusion, underscoring that alignment, not just more modalities, determines whether LLM-based music recommendation improves.

Original abstract

Music recommendation systems typically treat songs as opaque tokens, relying on collaborative interaction histories which overlooks semantic or acoustic content. Prior work has explored LLM-augmented, multimodal, and text-enhanced approaches to sequential recommendation, and while some methods partially combine semantic, acoustic, or engagement signals, none jointly model all three within a unified LLM-based sequential reasoning framework that grounds recommendations in actual song content. In this work, we propose a multimodal framework for session-based music recommendation that enriches the LastFM-1K dataset with three complementary signals: (1) audio and lyric embeddings extracted using pretrained music and text representation models, (2) LLM-generated semantic metadata using the MGPHot annotation schema, and (3) listening completion ratios. We adopt the E4SRec framework by extending it with multimodal features and different item ID encoder backbones, including SASRec, BERT4Rec, and GRU4Rec. We further extend the LLM backbone option with LLaMa-2-13B, Qwen2.5-7B-Instruct, and LLaMa-3-70B in both zero-shot and fine-tuned settings. Our experiments show that integrating content-based features improves over ID-only baselines up to 95% in terms of Recall and 79% in terms of NDCG. Moreover, our experiments show that naive multimodal fusion does not always yield additive improvements, highlighting challenges in cross-modal integration. We release a large-scale multimodal benchmark for music recommendation.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis