NTH

MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators

AuthorsYushi Huang, Xiangxin Zhou, Jun Zhang, Liefeng Bo, Tianyu Pang

July 18, 2026 2 min read
Watch on YouTube
The one-line take

MeanFlowNFT brings reinforcement-learning-based alignment to ultra-fast few-step image and video generators, achieving quality that can rival or exceed much slower diffusion models.

Key results

6
SD3.5-M metrics won

Best among few-step methods on 6 of 8 reported metrics.

1.4504
ImageReward

MeanFlowNFT score on SD3.5-M at 4 sampling steps.

0.6534
OCR

MeanFlowNFT OCR score on SD3.5-M at 4 sampling steps.

10
Function-evaluation reduction

10-fold fewer function evaluations than 40-step DiffusionNFT while matching or exceeding several metrics.

84.33
Wan2.1 VBench

MeanFlowNFT score at 4 steps on Wan2.1 1.3B.

82.57
LongCat-Video RL VBench

50-step comparison score surpassed by MeanFlowNFT.

What the paper found

MeanFlowNFT, from Tencent Hunyuan and The Hong Kong University of Science and Technology, is the first forward-process reinforcement-learning method designed for MeanFlow generators, which predict average velocity over time intervals rather than instantaneous velocity. Its key innovation is to derive an induced instantaneous-velocity predictor from the MeanFlow identity, apply the likelihood-free DiffusionNFT objective to that predictor, and retain MeanFlow’s native few-step sampler for inference. Theoretically, the authors show that, under idealized conditions, the induced optimum recovers the reward-improved policy and transfers that improvement back to the average-velocity generator. Using finite-difference derivatives, a shared reference derivative, and the forward conditional velocity for stable updates, MeanFlowNFT is evaluated on Stable Diffusion 3.5-Medium, or SD3.5-M, and Wan2.1 1.3B. On SD3.5-M, the 4-step model achieves the best result on 6 of 8 metrics, including an ImageReward of 1.4504 and OCR score of 0.6534, while matching or exceeding several metrics from 40-step DiffusionNFT with 10-fold fewer function evaluations. On video, 4-step MeanFlowNFT reaches a VBench score of 84.33 on Wan2.1 1.3B, surpassing the 50-step LongCat-Video RL score of 82.57, demonstrating that forward-process RL can improve quality without sacrificing few-step generation efficiency.

Original abstract

MeanFlow generators achieve fast few-step sampling by predicting average velocities over time intervals, making them attractive for efficient generation. Reinforcement learning (RL) has become a powerful way to align diffusion and flow models with human preferences and task-specific objectives. In particular, DiffusionNFT offers an efficient forward-process RL framework that does not require reverse-process trajectories or likelihood estimation. However, applying such RL methods to MeanFlow remains underexplored. DiffusionNFT optimizes instantaneous velocities, whereas MeanFlow samples with average velocities. To bridge this gap, we introduce MeanFlowNFT. Inspired by the MeanFlow identity, which bridges average and instantaneous velocities, we construct an induced instantaneous-velocity predictor. We apply the DiffusionNFT objective to this predictor, making reward optimization well-defined for MeanFlow. Sampling remains based on the average velocity, preserving MeanFlow's fast few-step generation. We further prove that MeanFlowNFT inherits DiffusionNFT's strict policy-improvement guarantee. Experiments on image and video generation show that MeanFlowNFT consistently improves baselines. Moreover, it outperforms prior state-of-the-art RL-tuned few-step generators on most metrics ($6$ of $8$ on SD3.5-M), and can even surpass multi-step RL-tuned diffusion while using only a few sampling steps. For instance, on Wan 2.1, $4$-step MeanFlowNFT reaches a VBench score of $84.33$, surpassing $50$-step LongCat-Video RL ($82.57$).

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis