Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models
AuthorsRuchit Rawal, Reza Shirkavand, Sayak Paul, Yuxin Wen, Heng Huang, Yizheng Chen, Tom Goldstein, Gowthami Somepalli
Resources
Flash-BoN speeds up diffusion model sampling by generating many cheap draft images first, then verifying and refining only the best one to get better results under the same compute budget.
Key results
Flash-BoN's improvement in area under the score-versus-wall-clock curve at larger model scales.
Speedup over BoN at matching its 300-second quality.
AUC/Time achieved after applying Flash drafts, versus 0.46 for Reflection-Tuning.
Pearson correlation between candidate-pool diversity and final GenAI-Bench performance.
Approximate reduction factor in gradient steps needed by Flash-Flow-GRPO to match baseline convergence.
What the paper found
Flash-BoN, from researchers at the University of Maryland and Hugging Face, challenges the assumption that diffusion-model inference-time scaling should spend computation on frequent intermediate verification. The paper shows that wall-clock measurement, including verifier overhead, can favor broad Best-of-N exploration over methods such as BFS, DFS, and ZOS. Its proposed pipeline first generates many inexpensive drafts using a per-model configuration combining timestep truncation, layer skipping, and second-order activation proxies; Qwen2.5-VL-7B then applies pointwise pruning followed by targeted pairwise Elo comparisons, and only the selected draft receives full-quality refinement. Across GenAI-Bench, GenEval, and UniGenBench, using Wan2.1 1.3B, Wan2.1 14B, and Black Forest Labs’ FLUX.1-dev, Flash-BoN leads under fixed wall-clock budgets, with gains reaching 8% AUC at larger model scales. On Wan2.1 1.3B, the full method achieves a 3.33x iso-quality speedup over BoN, while its combination with Reflection-Tuning raises AUC/Time from 0.46 to 0.62. Candidate diversity, measured with the Vendi Score over DINOv2 features, correlates with final quality at Pearson r=0.75. The same draft-guided selection idea also accelerates Flow-GRPO post-training: Flash-Flow-GRPO reaches baseline performance in roughly 10x fewer gradient steps, using 16 cheap draft rollouts to select 8 full trajectories.
Original abstract
Inference-time scaling for text-to-image generation has progressed from simple Best-of-$N$ (BoN) sampling to guided search methods that verify and steer candidate trajectories at intermediate denoising steps. These approaches focus on when and how often to verify during denoising but largely treat the cost of generation itself as fixed. Moreover, the standard practice of comparing methods by number of function evaluations (NFEs) counts only denoising forward passes and ignores verifier overhead, which can distort efficiency rankings. We show that under wall-clock evaluation, simple BoN already matches or outperforms several guided search techniques, suggesting that compute is better spent on broader exploration than on repeated intermediate verification. This motivates Flash-BoN, which generates a large pool of inexpensive draft candidates by combining three complementary acceleration knobs: timestep truncation, layer skipping, and activation proxies into a single configuration optimized once per model. An efficient multi-stage verification procedure then identifies the most promising draft, which is refined at full quality. Across three benchmarks and three model scales, Flash-BoN consistently outperforms all baselines under fixed wall-clock budgets, with gains that grow at larger model scales (+8% AUC). We further show that our strategy combines well and improves existing orthogonal techniques such as reflection-based prompt optimization (+16% AUC). The gains correlate with increased candidate diversity, which also enables draft-guided selection to accelerate RL post-training convergence.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.