Multi-Resolution Flow Matching: Training-Free Diffusion Acceleration via Staged Sampling
AuthorsXingyu Zheng, Xianglong Liu, Yifu Ding, Weilun Feng, Junqing Lin, Jinyang Guo, Haotong Qin
Resources
MrFlow speeds up text-to-image diffusion by generating a coarse image first and then refining it in stages, reaching large speedups without extra training.
Key results
MrFlow with 12 low-resolution steps and 1 high-resolution step on FLUX.1-dev at 1024×1024.
MrFlow with (12,1)×2 on Qwen-Image at 1024×1024.
MrFlow keeps OneIG-Bench within about 1% of native inference.
MrFlow combined with Pi-Flow on Qwen-Image.
Low-pass trajectory length relative to the full ODE path in the stage-wise analysis.
What the paper found
MrFlow, from Beihang University, ETH Zürich, and collaborators, is a training-free acceleration method for modern flow-matching diffusion models such as FLUX.1-dev and Qwen-Image. Instead of shrinking inference purely in latent space, it uses a staged low-to-high-resolution pipeline: generate the global structure at low resolution, decode to pixels, apply Real-ESRGAN super-resolution, re-encode, inject low-strength noise, then perform one high-resolution refinement step. The key idea is that low resolution cuts token cost by 4× per step and converges in fewer steps because the model spends about 58% of the original trajectory on low-frequency structure, while the high-resolution correction segment is close enough to the clean endpoint that a single Euler step is usually sufficient. On 1024×1024 evaluation, MrFlow reaches 8.25× speedup on FLUX.1-dev with 12+1 steps and 10.3× on Qwen-Image with (12,1)×2, while keeping OneIG-Bench within roughly 1% of native inference. It also composes cleanly with pretrained timestep distillation: MrFlow plus Pi-Flow reaches 11.3× on FLUX.1-dev and 25.1× on Qwen-Image. The paper argues that pixel-space GAN super-resolution is critical, because latent-space upsampling causes artifacts and blur, whereas low-strength noise selectively resamples high-frequency errors without erasing the preserved low-frequency layout.
Original abstract
Hardware-agnostic strategies for accelerating text-to-image diffusion, such as timestep distillation and feature caching, can reduce inference time without custom kernels or system-level optimization. Among them, multi-resolution generation strategies have recently received broad attention, attaining more than 5x speedup without any training. However, the design of performing upsampling in the latent space, together with the selective modification of partial regions, causes these methods to exhibit noticeable blurring or artifacts. To this end, we propose MrFlow, a training-free multi-resolution acceleration strategy for pretrained flow-matching models built upon a staged low-to-high-resolution pipeline. MrFlow first rapidly generates the main structure at low resolution, then performs super-resolution in the pixel space using a lightweight pretrained GAN-based model, subsequently injects low-strength noise to enable high-frequency resampling, and finally refines the details at high resolution. Quantitative and qualitative results on FLUX.1-dev and Qwen-Image show that MrFlow exploits the quadratic token reduction and reduced step requirement of low-resolution sampling to achieve 10x end-to-end acceleration while keeping OneIG within a 1% gap relative to that before acceleration, significantly surpassing other training-free acceleration strategies, and requiring no training or runtime dynamic identification whatsoever. MrFlow can further be directly combined orthogonally with pre-trained timestep distillation strategies, achieving even higher generation acceleration of up to 25x.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.