NTH
AI research

Stable Audio 3

AuthorsZach Evans, Julian D. Parker, Matthew Rice, CJ Carr, Zack Zukowski, Josiah Taylor, Jordi Pons

May 24, 2026 2 min read
Watch on YouTube
The one-line take

Stable Audio 3 is a fast, consumer-friendly audio generation and editing system that can create or modify music and sounds in seconds, even for variable-length clips.

Key results

6m 20s
max generation length

Stable Audio 3 supports variable-length stereo 44.1 kHz audio generation up to 6 minutes 20 seconds.

4096×
autoencoder downsampling

The semantic-acoustic autoencoder compresses stereo audio by 4096× before diffusion.

256-dimensional latents at approximately 10.76 Hz
latent size / rate

SAME produces compact 256-dimensional latent sequences for 44.1 kHz stereo audio.

0.100
Song Describer Dataset FAD

On the Song Describer Dataset at 190s, large achieves FAD 0.100, outperforming Stable Audio 2.5 at 0.128.

0.358
BBC Sound Effects FAD

On the BBC Sound Effects Dataset at 5s, large achieves FAD 0.358, outperforming Woosh Flow at 0.580.

What the paper found

Stable Audio 3 is a family of latent diffusion models for text-to-audio generation and editing that combines a 4096× semantic-acoustic autoencoder with a diffusion transformer, enabling stereo 44.1 kHz audio generation of up to 6 minutes 20 seconds. Its main novelty is native variable-length diffusion: instead of always generating a fixed maximum-length latent and wasting compute on silence, the model allocates latent length proportional to the requested duration, which makes short clips efficient and allows inference to scale from 0.44 seconds for 2-minute small models to 1.80 seconds for 6-minute large models on an H200 GPU. The autoencoder, based on SAME, preserves fidelity with multi-resolution STFT losses, adversarial training, chroma and interaural level difference regression, and contrastive latent alignment, while producing 256-dimensional latents at about 10.76 Hz. The diffusion transformer adds 64 learned memory embeddings, adaptive layer normalization for timestep and duration, T5Gemma text cross-attention, and inpainting through local-additive conditioning; medium and large further use differential attention. Training proceeds in three stages: flow matching with minibatch optimal transport coupling, a distillation warmup from a frozen teacher, and adversarial post-training using a relativistic discriminator plus CLAP-based text alignment. On the Song Describer Dataset, large achieves FAD 0.100 at 190 seconds versus 0.128 for Stable Audio 2.5, while on BBC Sound Effects it reaches FAD 0.358 at 5 seconds, outperforming Woosh Flow at 0.580. The open-weight small and medium checkpoints are released, with medium fitting in about 6.5 GB VRAM and small running on a MacBook Pro M4.

Original abstract

Stable Audio 3 is a family of fast latent diffusion models (small, medium, large) for variable-length audio generation and editing. Since our models can generate several minutes of audio, variable-length generations are key to avoid the cost of producing full-length generations for short sounds. We also support inpainting, enabling targeted audio editing and the continuation of short recordings. Our latent diffusion models operate on top of a novel semantic-acoustic autoencoder that projects audio into a compact latent space, enabling efficient diffusion-based generation while preserving audio fidelity and encouraging semantic structure in the latent. Finally, we run adversarial post-training to both accelerate inference and improve generation quality, reducing the number of inference steps while improving fidelity and prompt adherence. Stable Audio 3 models are trained on licensed and Creative Commons data to generate music and sounds in less than a 2s on an H200 GPU and less than a few seconds on a MacBook Pro M4. We release the weights of small and medium, that can run on consumer-grade hardware, together with their training and inference pipeline.

Read the original paper