NTH

Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models

AuthorsRuchit Rawal, Reza Shirkavand, Sayak Paul, Yuxin Wen, Heng Huang, Yizheng Chen, Tom Goldstein, Gowthami Somepalli

July 18, 2026 2 min read
Watch on YouTube
The one-line take

Flash-BoN speeds up diffusion model sampling by generating many cheap draft images first, then verifying and refining only the best one to get better results under the same compute budget.

Key results

8%
Large-scale AUC gain

Flash-BoN's improvement in area under the score-versus-wall-clock curve at larger model scales.

3.33
Wan2.1 1.3B iso-quality speedup

Speedup over BoN at matching its 300-second quality.

0.62
Flash-Reflection-Tuning AUC/Time

AUC/Time achieved after applying Flash drafts, versus 0.46 for Reflection-Tuning.

0.75
Diversity-quality correlation

Pearson correlation between candidate-pool diversity and final GenAI-Bench performance.

10
RL convergence acceleration

Approximate reduction factor in gradient steps needed by Flash-Flow-GRPO to match baseline convergence.

What the paper found

Flash-BoN, from researchers at the University of Maryland and Hugging Face, challenges the assumption that diffusion-model inference-time scaling should spend computation on frequent intermediate verification. The paper shows that wall-clock measurement, including verifier overhead, can favor broad Best-of-N exploration over methods such as BFS, DFS, and ZOS. Its proposed pipeline first generates many inexpensive drafts using a per-model configuration combining timestep truncation, layer skipping, and second-order activation proxies; Qwen2.5-VL-7B then applies pointwise pruning followed by targeted pairwise Elo comparisons, and only the selected draft receives full-quality refinement. Across GenAI-Bench, GenEval, and UniGenBench, using Wan2.1 1.3B, Wan2.1 14B, and Black Forest Labs’ FLUX.1-dev, Flash-BoN leads under fixed wall-clock budgets, with gains reaching 8% AUC at larger model scales. On Wan2.1 1.3B, the full method achieves a 3.33x iso-quality speedup over BoN, while its combination with Reflection-Tuning raises AUC/Time from 0.46 to 0.62. Candidate diversity, measured with the Vendi Score over DINOv2 features, correlates with final quality at Pearson r=0.75. The same draft-guided selection idea also accelerates Flow-GRPO post-training: Flash-Flow-GRPO reaches baseline performance in roughly 10x fewer gradient steps, using 16 cheap draft rollouts to select 8 full trajectories.

Original abstract

Inference-time scaling for text-to-image generation has progressed from simple Best-of-$N$ (BoN) sampling to guided search methods that verify and steer candidate trajectories at intermediate denoising steps. These approaches focus on when and how often to verify during denoising but largely treat the cost of generation itself as fixed. Moreover, the standard practice of comparing methods by number of function evaluations (NFEs) counts only denoising forward passes and ignores verifier overhead, which can distort efficiency rankings. We show that under wall-clock evaluation, simple BoN already matches or outperforms several guided search techniques, suggesting that compute is better spent on broader exploration than on repeated intermediate verification. This motivates Flash-BoN, which generates a large pool of inexpensive draft candidates by combining three complementary acceleration knobs: timestep truncation, layer skipping, and activation proxies into a single configuration optimized once per model. An efficient multi-stage verification procedure then identifies the most promising draft, which is refined at full quality. Across three benchmarks and three model scales, Flash-BoN consistently outperforms all baselines under fixed wall-clock budgets, with gains that grow at larger model scales (+8% AUC). We further show that our strategy combines well and improves existing orthogonal techniques such as reflection-based prompt optimization (+16% AUC). The gains correlate with increased candidate diversity, which also enables draft-guided selection to accelerate RL post-training convergence.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis