NTH
Research collection

Diffusion Models research

Follow diffusion research in image and video generation, sampling, and controllability. Compare quality improvements with their computational costs.

58 papers · Latest edition October 1, 2026

Where to start

Three of the latest briefs in this collection. Read the evidence and the original papers alongside them.

All Diffusion Models papers

Newest editions first.

02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis
07Diffusion

LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes

Chuyan Chen, Haoxing Chen, Kun Chen, Zhenglin Cheng, Long Cui, Ruishan Fang, Zhangxuan Gu, Zhicheng Huang, Zhenzhong Lan, Yuanting Lei, Haoquan Li, Jianguo Li, Rongchuan Li, Sidu Li, Tao Lin, Deyuan Liu, Jiacheng Liu, Lin Liu, Yuxuan Lou, Zhisheng Lu, Yuxin Ma, Shuheng Shen, Peng Sun, Chaoyang Wang, Hongjun Wang, Xiaomei Wang, Yongxin Wang, Chengzhang Wu, Hongru Wu, Jun Xie

LLaDA-Image offers an open, diffusion-based recipe for building highly capable image generators and distills them into a model that can generate and edit images in just a few steps.

Read analysis
09Diffusion

Unlocking Lossless Speedups in LLMs via Discrete Diffusion

Subham Sekhar Sahoo, Lingjie Chen, Khiem Pham, Jonathan Geuter, Chaitanya Dwivedi, Varad Pimpalkhute, Yash Akhauri, Alexander Moreno, Mikhail Yurochkin, Zhenting Wang, Mostafa Elhoushi, Nolan Dey, Shane Bergsma, Joel Hestness, John Thickstun, Eric Xing, Zhengzhong Liu

Uno uses discrete diffusion to make standard autoregressive LLMs generate multiple tokens in parallel, delivering up to three times faster inference without sacrificing quality.

Read analysis
10Diffusion

4DStreamCtrl: Interactive Video Generation with Online 4D Control

Shiqian Li, Chenguo Lin, Zhiguang Liu, Yu Tang, Jiarong Ou, Rui Chen, Yixin Zhu

4DStreamCtrl aims to make video diffusion interactive by generating long, realistic streams while jointly controlling camera motion, object movement, and depth in real time.

Read analysis
12Diffusion

Pixel-Space Diffusion via Observation Operators

Shaojie Guo, Lichen Ma, Haoyang Tong, Yu He, Zipeng Guo, Xiaoan Liu, Feng Yan, Yu Guo, Fei Wang, Junshi Huang, Yan Wang

This work makes pixel-space diffusion easier to train by teaching models to recover images from coarse structures to fine details in the same order humans and vision systems naturally perceive them.

Read analysis
13Diffusion

Abra: Scaling Diffusion Image Training

Kyle Chickering, Wei-An Lin, Swayam Bhanded, Dan Saunders, Akshat Tripathi, Jiaming Song, Shyamal Buch, Xinchen Yan

Abra shows that diffusion image models scale predictably, but reach their best performance with far more training data per parameter than language models.

Read analysis
14Diffusion

An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models

Dengyang Jiang, Ruoyi Du, Zhennan Chen, Dongyang Liu, Zanyi Wang, Mingzhe Zheng, Xiangpeng Yang, Huanqia Cai, Aiming Hao, Yuming Jiang, Peng Gao, Harry Yang, Steven Hoi

This study shows how to efficiently train pixel-space image diffusion models by first learning in latent space, achieving competitive quality with substantially faster inference.

Read analysis
15Diffusion

LinCa: Accelerating Diffusion Models via Learnable Decomposed Feature Caching

Jinshan Liu, Haoran Qin, Xiaobing Tu, Jiacheng Liu, Jiahui Hu, Zhengan Yan, Yukun Xie, Kerui Shen, Jinkui Ren, Yuqi Lin, Xiantao Zhang, Linfeng Zhang

LinCa speeds up diffusion-based image and video generation by learning which parts of intermediate features can be safely reused, achieving substantial acceleration with minimal quality loss.

Read analysis
16Diffusion

DiffusionGemma Technical Report

DiffusionGemma Team, Adrien Ali Taïga, James Assiene, Daniele Calandriello, Rahma Chaabouni, João Gante, Tamara von Glehn, Nate Keating, Chris Knutsen, Martin Kukla, Tianlin Liu, Ivan Lobov, Ofir Nabati, João Gabriel Oliveira, Nicolas Perez-Nieves, Nastasia Prutianova, Bobak Shahriari, Jean Tarbouriech, Pavel Tyletski, Çağlar Ünlü, Cindy Wu, Glenn Cameron, Jerome Connor, Sertan Girgin, Maarten Grootendorst, Alon Levkovitch, Eliya Nachmani, Omar Sanseviero, Piotr Stanczyk, Quentin Berthet, Andrew Campbell, Clément Crepy, Valentin De Bortoli, Arnaud Doucet, Romuald Elie, Alexandre Galashov, Klaus Greff, Alexis Jacq, David Ruhe, Yu-Han Wu, Sebastian Flennerhag, Brendan O'Donoghue, George Scrivener, Shantanu Thakoor

DiffusionGemma turns a large language model into a fast parallel text generator, reaching roughly 1,500 tokens per second while preserving much of the model's reasoning and multimodal capability.

Read analysis
17Diffusion

JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion

Yicheng Xiao, Wenxun Dai, Xinran Qin, Lin Song, Maoquan Zhang, Hang Xu, Yukang Chen, Yitong Li, Guohui Zhang, Yuan Zhang, Xuying Zhang, Tommy Zhang, Jianlong Yuan, Peihao Li, Shuai Lu, Siming Fu, Chuyang Zhao, Xin Han, Jie Huang, Wenbo Li, Guoqing Ma, Wei Huang, Xiaojuan Qi, Haoyang Huang, Nan Duan

JoyAI-Video-Edit aims to make high-quality, open-ended video editing run in real time by combining autoregressive generation with diffusion distillation.

Read analysis
18Diffusion

Scaling Properties of Text Conditioning in Visual Generation

Zilong Chen, Chaorui Deng, Kunchang Li, Hongyi Yuan, Haoqi Fan

This work shows that better-structured language can make visual diffusion models more capable, and uses scaling laws to build a stronger prompting system.

Read analysis
24Diffusion

Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers

Maohua Li, Qirui Li, Yanke Zhou, Yiduo Li, Zhaosheng Chi, Chao Xu, Cuifeng Shen, Yixuan Xu, Hanlin Tang, Kan Liu, Tao Lan, Lin Qu, Shao-Qun Zhang

This work shows that seemingly uninformative template tokens in diffusion transformers quietly preserve object identity, and that some attention heads can be pruned to save computation with modest quality loss.

Read analysis
25Diffusion

Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models

Ruchit Rawal, Reza Shirkavand, Sayak Paul, Yuxin Wen, Heng Huang, Yizheng Chen, Tom Goldstein, Gowthami Somepalli

Flash-BoN speeds up diffusion model sampling by generating many cheap draft images first, then verifying and refining only the best one to get better results under the same compute budget.

Read analysis
26Diffusion

Hierarchical Denoising For Multi-Step Visual Reasoning

Zezhong Qian, Xiaowei Chi, Chak-Wing Mak, Tianze Zhou, Ruibin Yuan, Yuhan Rui, Hengzhe Sun, Zhuoqun Wu, Yuming Li, Siyuan Qian, Sirui Han, Shanghang Zhang

HDR lets video diffusion models plan complex visual actions hierarchically before streaming them efficiently, improving reasoning while sharply reducing inference cost.

Read analysis
28Diffusion

Cross-Space Distillation: Teaching One-Step Students with Modern Diffusion Teachers

Anh Nguyen, Ngan Nguyen, Duc Vu, Trung Dao, Viet Nguyen, Quan Dao, Kien Nguyen, Chi Tran, Phong Nguyen, Khoi Nguyen, Cuong Pham, Dimitris Metaxas, Vishal M. Patel, Anh Tran

This paper shows how to distill powerful modern diffusion models into smaller one-step generators even when the teacher and student use incompatible latent spaces.

Read analysis
34Diffusion

Exploring the Design Space of Reward Backpropagation for Flow Matching

Ruoyu Wang, Boye Niu, Xiangxin Zhou, Yushi Huang, Tongliang Liu, Chi Zhang

This paper makes it cheaper and more stable to train image generators from human preferences by redesigning how reward gradients are backpropagated through flow-matching sampling trajectories.

Read analysis
35Diffusion

Safe Few-Step Generation via Velocity Editing

Yujin Choi, Jaehong Yoon

This paper introduces a training-free way to make fast text-to-image generators safer by editing their velocity trajectories, cutting down harmful outputs while keeping benign images intact.

Read analysis
37Diffusion

Colored Noise Diffusion Sampling

Hadar Davidson, Noam Issachar, Sagie Benaim

This paper improves diffusion image generation by replacing uniform noise with frequency-aware colored noise during sampling, making the process more efficient and producing sharper images.

Read analysis
39Diffusion

JLT: Clean-Latent Prediction in Latent Diffusion Transformers

Funing Fu, Tenghui Wang, Junyong Cen, Qichao Zhu, Guanyu Zhou

This paper shows that in latent diffusion models, predicting the clean latent directly can work better than predicting velocity, because the choice of training target changes the geometry of learning in compressed latent space.

Read analysis
40Diffusion

PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion

Yifan Lu, Qi Wu, Jay Zhangjie Wu, Zian Wang, Huan Ling, Sanja Fidler, Xuanchi Ren

PiD replaces traditional latent decoders with a fast pixel-diffusion decoder that can turn compressed latents into high-resolution images much more efficiently and with better visual fidelity.

Read analysis
41Diffusion

Score-Control for Hallucination Reduction in Diffusion Models

Mahesh Bhosale, Naresh Kumar Devulapally, Abdul Wasi, Chau Pham, Vishnu Suresh Lokhande, David Doermann

This paper tackles hallucinations in diffusion image generators by controlling the model’s score dynamics, improving reliability without sacrificing quality.

Read analysis
44Diffusion

Efficient and Training-Free Single-Image Diffusion Models

Haojun Qiu, Kiriakos N. Kutulakos, David B. Lindell

This paper turns single-image diffusion into a fast, training-free process by using closed-form patch statistics, enabling high-quality image generation from one reference image in seconds rather than hours.

Read analysis
46Diffusion

High-Fidelity Two-Step Image Generation via Teacher-Aligned End-to-End Distillation

Dongyang Liu, Ruoyi Du, David Liu, Dengyang Jiang, Liangchen Li, Qilong Wu, Zhen Li, Steven C. H. Hoi, Hongsheng Li, Peng Gao

This paper shows how to squeeze high-quality images out of a diffusion model in just two denoising steps by aligning the student to its teacher, separating step-specific parameters, and training end-to-end with regularization.

Read analysis
48Diffusion

i1: A Simple and Fully Open Recipe for Strong Text-to-Image Models

Boya Zeng, Tianze Luo, Shu Pu, Jucheng Shen, Taiming Lu, Gabriel Sarch, Zhuang Liu

This paper delivers a fully open, highly competitive text-to-image diffusion model plus a carefully tested training recipe that could become a new starting point for open generative AI research.

Read analysis
49Diffusion

Injecting Image Guidance into Text-Conditioned Diffusion Models at Inference

Agata Żywot, Iason Skylitsis, Thijmen Nijdam, Zoe Tzifa-Kratira, Derck Prinzhorn, Konrad Szewczyk, Aritra Bhowmik

This paper lets Stable Diffusion take both a text prompt and a reference image at inference time, so you can steer generation toward a desired style or visual concept without retraining.

Read analysis
51Diffusion

Rethinking Cross-Layer Information Routing in Diffusion Transformers

Chao Xu, Maohua Li, Qirui Li, Yixuan Xu, Yanke Zhou, Yunhe Li, Cuifeng Shen, Hanlin Tang, Kan Liu, Tao Lan, Lin Qu, Shao-Qun Zhang

This paper argues that the way diffusion transformers pass information across layers is suboptimal, and introduces a new residual-routing method that can speed up training and improve image quality.

Read analysis