Retrieving Individual Stems from Music Mixtures with Slot Embeddings
AuthorsDavid Braun, Junyi Fan, Pranay Manocha, Donald S. Williamson, Adam Finkelstein
Resources
Stembed lets music producers search for individual instrument sounds hidden inside a full song by representing the mixture as multiple searchable stem-like embeddings.
Key results
Stembed’s rank-one recall across the 2068-stem gallery.
Pooled mixture-embedding baseline on the same MoisesDB task.
Stembed’s rank-one recall when the target instrument family is supplied.
Recall when Stembed predicts the number of active stems rather than receiving it.
Stembed’s rank-one recall on the independent Slakh2100 benchmark.
Total parameters in the selected Stembed model.
What the paper found
Stembed is a music-retrieval model that decomposes a mixture into four slot embeddings, each intended to represent a candidate stem, instead of compressing the entire mix into one vector. It uses the MuQ self-supervised music encoder, a transformer slot decoder, minimum-cost matching between mixture slots and isolated stems, supervised contrastive learning with stem identity labels, and a presence predictor that selects usable slots at inference. Training constructs two- and three-stem mixtures from the same song, while evaluation uses MoisesDB and the independent Slakh2100 benchmark. On three-stem MoisesDB retrieval across the full 2068-stem gallery, Stembed achieves 56.8% R@1, versus 23.6% for the pooled CIR–MuQ baseline; with instrument-family filtering, it reaches 58.9% R@1 versus CIR–MuQ’s 52.4%. Without supplying the stem count, its presence threshold still produces 56.0% R@1. On Slakh2100, Stembed reaches 38.1% full-gallery R@1, compared with 20.2% for CIR–MuQ, indicating transfer beyond the training corpus. Direct mixture retrieval also outperforms HT-Demucs separation followed by Stembed, which achieves 53.2% R@1. The selected system has 191M parameters, outputs 256-dimensional normalized retrieval embeddings, and demonstrates that slot specialization can replace explicit instrument-family labels with a recognition-and-selection workflow.
Original abstract
Music producers search libraries of isolated instrument recordings, called stems, for sounds resembling parts of an existing song. Neural retrieval systems address this by mapping audio to embeddings and ranking library stems by their similarity to the query. The leading method, Contrastive Instrument Retrieval (CIR), encodes the mixture as a single embedding, but it works best when a user specifies the target's instrument family. We introduce Stembed, which encodes a mixture as several slot embeddings representing candidate stems. During training, we construct mixtures from stems of the same song and match their slot embeddings to those of the isolated stems. The slot embeddings from mixtures inherit the stem identities of their assigned solo embedding, enabling a contrastive loss. On mixtures from held out MoisesDB artists, Stembed outperforms a CIR-style baseline when both search the full stem library. Even when predicting the stem count itself without family labels, Stembed exceeds the baseline's family-filtered R@1. Our website demonstrates how users can select a slot by inspecting the tags of its retrieved stems.
Read the original paperMore in Neural Networks
Browse all 22 papers →End-to-End Hard-Label Cryptanalytic Model Extraction Using Efficient Sign Recovery
Akira Ito, Takayuki Miura, Yosuke Todo
A new query-efficient technique makes it possible to steal the parameters of small black-box neural networks using only their predicted labels.
The Linear Representation Hypothesis Needs a Group Action
Louie Hong Yao, Yuhao Li, Shengchao Liu
This paper argues that claims about linear representations only become meaningful once we specify which transformations leave a representation essentially unchanged.
HypLTSF: A Hyperbolic Geometric View of Multi-Scale Hierarchies for Long-Term Time Series Forecasting
Namwoo Kim, Hyungryul Baik, Yoonjin Yoon
HypLTSF maps multi-scale time-series patterns into hyperbolic space so that fine-to-coarse temporal hierarchies become explicit and useful for long-term forecasting.