DEMON: Diffusion Engine for Musical Orchestrated Noise
AuthorsRyan Fosdick
Resources
DEMON turns a diffusion music generator into a live, playable instrument by speeding up and restructuring how denoising controls take effect in real time.
Key results
Sustained throughput on a single RTX 5090 for 60-second music outputs at depth 8.
Measured throughput at the production ring depth of 4.
Continuous slider sweep baseline for StreamDiffusion-style queue reset.
Continuous slider sweep with heterogeneous per-slot scheduling.
Windowed VAE decode latency reduction from rendering only the playback window with overlap margins.
What the paper found
DEMON, from Daydream, turns ACE-Step 1.5, a diffusion music model, into a real-time performance instrument by adapting StreamDiffusion’s ring-buffer streaming design to audio and accelerating it with TensorRT. On an RTX 5090, it sustains 12.3 decoder completions per second for 60-second tracks, or 11.3 generations per second at the production depth of 4, while maintaining sample-identical parity with batch decoding at the 16-bit PCM level. Its main novelty is a latency taxonomy for controllable diffusion: per-request parameters such as scalar denoise still obey an S-tick drain floor, but DEMON introduces per-slot heterogeneous denoise scheduling so a moving slider does not wipe in-flight generations; shared mutable per-step state for controls like SDE source blending, guidance, and x0-target morphing, which takes effect on the very next denoising tick; and a windowed VAE decode that cuts decode latency up to 8.0x by rendering only the playback window with overlap margins. Experiments show the per-slot design preserves 100 percent completion under continuous slider sweeps, versus 1.7 percent for a StreamDiffusion-style queue reset, and the shared-step controls produce true per-frame gradients across 60-second outputs while preserving CLAP alignment and matching batch-mode audio exactly. The paper positions the system as a controllable, low-latency diffusion engine for live music, while noting that perceptual validation with musicians remains future work.
Original abstract
We present DEMON, a real-time diffusion engine that makes the denoising process playable as a live musical instrument: a control surface both broad (many parameters shaped per-frame across the output) and responsive (each control taking effect as fast as its place in the denoising loop allows). Built on ACE-Step 1.5 and StreamDiffusion's ring-buffer architecture with TensorRT acceleration, it sustains up to 12.3 decoder completions per second for 60-second music on a single consumer GPU (RTX 5090), or 11.3 generations per second at our production ring-depth of 4. At these rates denoising parameters become viable as live performance controls, but the ring buffer propagates per-request changes only at its drain rate, a floor of S denoising steps. We contribute four mechanisms. (1) Per-slot heterogeneous denoise scheduling: each ring-buffer slot owns its timestep schedule, so a moving denoise slider is tracked without wiping the in-flight queue, where the upstream global-schedule design must rebuild and discard it. (2) Shared mutable per-step state, giving any parameter consulted at every solver step next-tick effect, bypassing ring-buffer drain. (3) Per-frame source blending: a sampling-time control on the standard SDE re-noise step, giving a framewise transformation-strength axis that complements scalar denoise scheduling. (4) Windowed VAE decode exploiting receptive-field analysis for an 8.0x decode speedup. Together these separate streaming-diffusion parameters into four propagation classes, by onset and convergence latency.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.