NTH

Perceptual Flow Matching for Few-Step Generative Modeling

AuthorsChuyang Zhao, Yifei Song, Hongfa Wang, Jianlong Yuan, Yuan Zhang, Siming Fu, Zhineng Chen, Huilin Deng, Haoyang Huang, Nan Duan

July 8, 2026 2 min read
Watch on YouTube
The one-line take

This paper makes generative models faster by training them in perceptual feature space, letting them produce good samples in just a few steps instead of dozens.

Key results

33.93
COCO FID

PFM on SD3-Medium at 8 steps

31.70
COCO CLIP

PFM on SD3-Medium at 8 steps

11.42
COCO HPSv3

PFM on SD3-Medium at 8 steps

0.9402
MagicBrush CLIP-I

PFM on Qwen-Image-Edit at 8 steps

0.9187
MagicBrush DINO

PFM on Qwen-Image-Edit at 8 steps

What the paper found

Perceptual Flow Matching, from Joy Future Academy, Fudan University, Tsinghua University, and USTC, reframes few-step flow-matching generation by moving supervision out of VAE latent space and into pretrained perceptual feature spaces such as VGG, DINOv2, SigLIP, ConvNeXt, and InternVideo2. Instead of regressing velocity with Euclidean loss, the model decodes the predicted clean sample and optimizes a perceptual distance, which the paper argues shifts learning from mean-seeking toward mode-seeking and makes coarse integration far less prone to blur. This simple change requires no teacher model, auxiliary score network, or distillation pipeline, yet reduces sampling from 35–50 steps to 4–8 steps while preserving quality. On SD3-Medium for text-to-image generation, PFM reaches state-of-the-art few-step performance on COCO 2014 val with 8 steps, obtaining 33.93 FID, 31.70 CLIP, and 11.42 HPSv3. On MagicBrush image editing with Qwen-Image-Edit, it matches or surpasses the 40-step baseline using only 8 steps and no inference-time classifier-free guidance, including 0.9402 CLIP-I and 0.9187 DINO. On Wan2.1-1.3B video generation with InternVideo2 supervision, it improves VBench and slightly exceeds the 35-step baseline on the overall score, while remaining strong at 8 steps. Ablations show that pretrained perceptual backbones matter: RandViT fails, and stronger off-manifold penalties correlate with better few-step results.

Original abstract

We propose Perceptual Flow Matching (PFM), a simple yet effective framework for few-step generation in flow-matching models. Rather than performing velocity regression in the conventional VAE latent space, PFM supervises flow matching in a perceptual feature space using pretrained perceptual models. This simple change substantially improves the few-step generation capability of flow-matching models, reducing the number of sampling steps from 35-50 to 4-8 while preserving generation quality. Unlike existing acceleration and distillation approaches, PFM requires neither teacher models nor auxiliary score networks and can be integrated into standard flow-matching training pipelines with minimal modifications. Extensive experiments on image generation, video generation, and image editing tasks demonstrate that PFM consistently produces high-quality results while producing fewer artifacts than existing distillation-based methods. We further show that perceptual supervision shifts the regression minimizer from mean-seeking to mode-seeking, biasing predictions toward on-manifold modes that remain accurate under coarse few-step integration. Our results reveal that standard flow-matching training can naturally yield high-quality few-step generators when supervised in an appropriate representation space. We hope this insight inspires future research into representation-aware objectives for efficient generative modeling.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis