FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
AuthorsJaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
AffiliationsInterdisciplinary Program in AI, Seoul National University · Department of Electrical and Computer Engineering, Seoul National University · AIIS, ASRI, INMC, and ISRC, Seoul National University
Resources
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
Key results
Spearman correlation between diffusion fork ordering and human-perceived similarity
Individual human responses agreeing with the fork-induced ordering
Automatically generated reference–variant pairs used for training
Maximum FLUX.1-dev denoising steps used for sampling fork points
PIPAL SROCC achieved with the LPIPS-Alex backbone
Average PIPAL SROCC across seven evaluated backbones
What the paper found
FoMo, or Forking Moment, turns a diffusion model’s denoising trajectory into an automatic perceptual-distance label for reference-based image quality assessment. Starting from a reference image, the method injects noise at a selected point in the trajectory and independently denoises afterward: early forks discard shared structure and produce perceptually distant images, while late forks preserve detail and produce close variants. Using FLUX.1-dev with a 50-step schedule, FoMo generated 480k labeled image pairs without human annotation, combining ImageNet references with synthetic references. Instead of conventional MOS regression or 2AFC classification, it trains with a RankNet-style RankBCE objective that compares every pair in a batch, enforcing globally consistent ordering across unrelated reference images. Human validation found a Spearman correlation of 0.970 between fork ordering and perceived similarity, while a cross-reference forced-choice study matched the fork-induced ordering in 90.2% of responses. Across PIPAL, FoMo achieved a 0.733 SROCC using an LPIPS-Alex backbone, compared with 0.577 for KADID-10K supervision, and its average SROCC across seven backbones on PIPAL was 0.683. The approach also worked with Stable Diffusion 1.5, SD-XL, and SD3, although FLUX.1 was strongest. Its main limitation is coverage: because diffusion-generated examples do not naturally represent localized corruption, pixel noise, or purely photometric shifts, performance can weaken on those out-of-distribution distortions.
Original abstract
Reference-based image quality assessment (IQA) metrics aim to reflect how humans perceive the perceptual distance between a pair of images. To learn how the human visual system (HVS) operates, recent reference-based IQA metrics heavily rely on human-annotated data. Mean opinion score (MOS)-based pointwise scoring, which assigns a scalar quality value per image, is preferable for annotation but is prohibitively expensive to collect at scale and is known to be noisy due to inconsistent human judgments. As an alternative, two-alternative forced choice (2AFC) pairwise labels have gained popularity due to their reliability and efficiency, but they capture only relative comparisons between pairs. In this paper, we propose a fully automated data generation pipeline that generates pointwise perceptual distance labels between image pairs without any human annotation. Our approach exploits the generative dynamics of diffusion models as a perceptual distance proxy, where the coarse structure of an image is generated in the early timesteps and the fine details are generated in the later timesteps. Images that fork early in the generation process share only coarse structure and are perceptually far apart; images that fork late differ only in fine detail. We demonstrate that the diffusion trajectory aligns well with the human visual system, and use this forking moment, FoMo, as a reference-grounded distance label to supervise the training of a reference-based IQA metric. The pointwise labels, which support universal comparison between arbitrary image pairs, enable an information-rich training objective. Extensive experiments across diverse backbone architectures confirm the effectiveness of our generation pipeline, outperforming human-annotated datasets in multiple benchmarks.
Read the original paperMore in Diffusion Models
Browse all 58 papers →LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.
Register Tokens for Bounded-State Reasoning in Diffusion Language Models
Albert Ge, Chandan Singh, Yufan Zhuang, Xiaodong Liu, Jianfeng Gao, Frederic Sala
The work gives diffusion language models a small set of learned memory tokens so they can preserve reasoning across multiple chunks without retaining the generated text.