NTH

DEMON: Diffusion Engine for Musical Orchestrated Noise

AuthorsRyan Fosdick

June 7, 2026 2 min read
Watch on YouTube
The one-line take

DEMON turns a diffusion music generator into a live, playable instrument by speeding up and restructuring how denoising controls take effect in real time.

Key results

12.3
decoder completions per second

Sustained throughput on a single RTX 5090 for 60-second music outputs at depth 8.

11.3
production generations per second

Measured throughput at the production ring depth of 4.

1.7%
global-reset completion rate

Continuous slider sweep baseline for StreamDiffusion-style queue reset.

100%
per-slot sweep completion rate

Continuous slider sweep with heterogeneous per-slot scheduling.

8.0x
VAE decode speedup

Windowed VAE decode latency reduction from rendering only the playback window with overlap margins.

What the paper found

DEMON, from Daydream, turns ACE-Step 1.5, a diffusion music model, into a real-time performance instrument by adapting StreamDiffusion’s ring-buffer streaming design to audio and accelerating it with TensorRT. On an RTX 5090, it sustains 12.3 decoder completions per second for 60-second tracks, or 11.3 generations per second at the production depth of 4, while maintaining sample-identical parity with batch decoding at the 16-bit PCM level. Its main novelty is a latency taxonomy for controllable diffusion: per-request parameters such as scalar denoise still obey an S-tick drain floor, but DEMON introduces per-slot heterogeneous denoise scheduling so a moving slider does not wipe in-flight generations; shared mutable per-step state for controls like SDE source blending, guidance, and x0-target morphing, which takes effect on the very next denoising tick; and a windowed VAE decode that cuts decode latency up to 8.0x by rendering only the playback window with overlap margins. Experiments show the per-slot design preserves 100 percent completion under continuous slider sweeps, versus 1.7 percent for a StreamDiffusion-style queue reset, and the shared-step controls produce true per-frame gradients across 60-second outputs while preserving CLAP alignment and matching batch-mode audio exactly. The paper positions the system as a controllable, low-latency diffusion engine for live music, while noting that perceptual validation with musicians remains future work.

Original abstract

We present DEMON, a real-time diffusion engine that makes the denoising process playable as a live musical instrument: a control surface both broad (many parameters shaped per-frame across the output) and responsive (each control taking effect as fast as its place in the denoising loop allows). Built on ACE-Step 1.5 and StreamDiffusion's ring-buffer architecture with TensorRT acceleration, it sustains up to 12.3 decoder completions per second for 60-second music on a single consumer GPU (RTX 5090), or 11.3 generations per second at our production ring-depth of 4. At these rates denoising parameters become viable as live performance controls, but the ring buffer propagates per-request changes only at its drain rate, a floor of S denoising steps. We contribute four mechanisms. (1) Per-slot heterogeneous denoise scheduling: each ring-buffer slot owns its timestep schedule, so a moving denoise slider is tracked without wiping the in-flight queue, where the upstream global-schedule design must rebuild and discard it. (2) Shared mutable per-step state, giving any parameter consulted at every solver step next-tick effect, bypassing ring-buffer drain. (3) Per-frame source blending: a sampling-time control on the standard SDE re-noise step, giving a framewise transformation-strength axis that complements scalar denoise scheduling. (4) Windowed VAE decode exploiting receptive-field analysis for an 8.0x decode speedup. Together these separate streaming-diffusion parameters into four propagation classes, by onset and convergence latency.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis