Inference-Time Scaling of Diffusion Models via Progressive Seed Pruning
AuthorsRogerio Guimaraes, Pietro Perona
The method improves image generation by exploring many noise seeds early, pruning weak candidates progressively, and spending computation only on the most promising ones.
Key results
PSP matches the compute of four full denoising trajectories while exploring a larger seed pool.
PSP’s GenEval score on GenEval prompts at matched compute.
PSP’s GenEval score for Stable Diffusion XL.
PSP’s GenEval score for the flow-matching Stable Diffusion 3.5 model.
Human prompt-alignment score for PSP on SDXL.
Seconds required by PSP at compute multiplier 4 on an H200.
What the paper found
Researchers Rogério Guimarães and Pietro Perona at Caltech introduce Progressive Seed Pruning, or PSP, a training-free, gradient-free method for scaling diffusion and flow-matching inference under a fixed denoising budget. Instead of fully denoising a constant number of samples, PSP starts with a larger pool of noise seeds, scores intermediate clean-image estimates using the black-box ImageReward model, and progressively retains only the most promising trajectories. Its default schedule uses an effective compute multiplier of 4, begins with 8 seeds, prunes to 4 at 25 percent progress, and then to 2 at 50 percent. On GenEval prompts, PSP reaches GenEval scores of 0.574 for Stable Diffusion v1.5, 0.645 for SDXL, and 0.747 for Stable Diffusion 3.5, outperforming Best-of-N, FK-Steering, and DSearch at matched compute; human prompt-alignment evaluation gives SDXL a score of 0.713. PSP remains effective with deterministic samplers, which is important for flow-matching models, and its intermediate rewards can be cached to search pruning schedules offline. On an H200, SDXL PSP at multiplier 4 takes 13.74 seconds, close to 12.48 seconds for Best-of-4 and substantially below 23.90 seconds for Best-of-8. The authors also extend PSP to prompt selection using 24 rewrites generated by OpenAI’s ChatGPT 5.2, suggesting the method can optimize any discrete candidate pool rather than only random seeds.
Original abstract
Diffusion and flow-matching models dominate conditional image generation, yet inference-time scaling for these models is far less developed than for autoregressive language models. Because final quality is highly sensitive to the initial noise seed, many approaches spend extra compute on seed search or resampling under a black-box reward, but typically maintaining a constant memory footprint throughout inference. We show that relaxing this constraint enables an underexplored inference-time scaling axis: by front-loading exploration, evaluating many seeds early, and pruning aggressively, we can use a fixed compute budget more effectively. \emph{Progressive Seed Pruning} (\PSP) scores intermediate denoised estimates and progressively narrows the candidate set so that only promising trajectories are fully denoised, while keeping the total number of model evaluations fixed. Across diffusion and flow-matching backbones, \PSP \ consistently improves reward-guided selection and achieves higher GenEval scores (automated) and better human evaluation on prompt-alignment than best-of-$N$, importance-sampling, and tree-search baselines at matched compute. Project page: https://www.vision.caltech.edu/psp. Code: https://github.com/rogerioagjr/psp.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.