NTH

Retrieving Individual Stems from Music Mixtures with Slot Embeddings

AuthorsDavid Braun, Junyi Fan, Pranay Manocha, Donald S. Williamson, Adam Finkelstein

September 29, 2026 2 min read
Watch on YouTube
The one-line take

Stembed lets music producers search for individual instrument sounds hidden inside a full song by representing the mixture as multiple searchable stem-like embeddings.

Key results

56.8%
MoisesDB full-gallery R@1

Stembed’s rank-one recall across the 2068-stem gallery.

23.6%
CIR–MuQ full-gallery R@1

Pooled mixture-embedding baseline on the same MoisesDB task.

58.9%
MoisesDB family-filtered R@1

Stembed’s rank-one recall when the target instrument family is supplied.

56.0%
Predicted-count R@1

Recall when Stembed predicts the number of active stems rather than receiving it.

38.1%
Slakh2100 full-gallery R@1

Stembed’s rank-one recall on the independent Slakh2100 benchmark.

191M
Model parameters

Total parameters in the selected Stembed model.

What the paper found

Stembed is a music-retrieval model that decomposes a mixture into four slot embeddings, each intended to represent a candidate stem, instead of compressing the entire mix into one vector. It uses the MuQ self-supervised music encoder, a transformer slot decoder, minimum-cost matching between mixture slots and isolated stems, supervised contrastive learning with stem identity labels, and a presence predictor that selects usable slots at inference. Training constructs two- and three-stem mixtures from the same song, while evaluation uses MoisesDB and the independent Slakh2100 benchmark. On three-stem MoisesDB retrieval across the full 2068-stem gallery, Stembed achieves 56.8% R@1, versus 23.6% for the pooled CIR–MuQ baseline; with instrument-family filtering, it reaches 58.9% R@1 versus CIR–MuQ’s 52.4%. Without supplying the stem count, its presence threshold still produces 56.0% R@1. On Slakh2100, Stembed reaches 38.1% full-gallery R@1, compared with 20.2% for CIR–MuQ, indicating transfer beyond the training corpus. Direct mixture retrieval also outperforms HT-Demucs separation followed by Stembed, which achieves 53.2% R@1. The selected system has 191M parameters, outputs 256-dimensional normalized retrieval embeddings, and demonstrates that slot specialization can replace explicit instrument-family labels with a recognition-and-selection workflow.

Original abstract

Music producers search libraries of isolated instrument recordings, called stems, for sounds resembling parts of an existing song. Neural retrieval systems address this by mapping audio to embeddings and ranking library stems by their similarity to the query. The leading method, Contrastive Instrument Retrieval (CIR), encodes the mixture as a single embedding, but it works best when a user specifies the target's instrument family. We introduce Stembed, which encodes a mixture as several slot embeddings representing candidate stems. During training, we construct mixtures from stems of the same song and match their slot embeddings to those of the isolated stems. The slot embeddings from mixtures inherit the stem identities of their assigned solo embedding, enabling a contrastive loss. On mixtures from held out MoisesDB artists, Stembed outperforms a CIR-style baseline when both search the full stem library. Even when predicting the stem count itself without family labels, Stembed exceeds the baseline's family-filtered R@1. Our website demonstrates how users can select a slot by inspecting the tags of its retrieved stems.

Read the original paper

More in Neural Networks

Browse all 22 papers →
02Neural Network

The Linear Representation Hypothesis Needs a Group Action

Louie Hong Yao, Yuhao Li, Shengchao Liu

This paper argues that claims about linear representations only become meaningful once we specify which transformations leave a representation essentially unchanged.

Read analysis