NTH

FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance

AuthorsJaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon

AffiliationsInterdisciplinary Program in AI, Seoul National University · Department of Electrical and Computer Engineering, Seoul National University · AIIS, ASRI, INMC, and ISRC, Seoul National University

October 1, 2026 3 min read
Watch on YouTube
The one-line take

FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.

Key results

0.970
Human alignment correlation

Spearman correlation between diffusion fork ordering and human-perceived similarity

90.2%
Cross-reference agreement

Individual human responses agreeing with the fork-induced ordering

480k
FoMo training pairs

Automatically generated reference–variant pairs used for training

50
Diffusion schedule

Maximum FLUX.1-dev denoising steps used for sampling fork points

0.733
PIPAL FoMo SROCC

PIPAL SROCC achieved with the LPIPS-Alex backbone

0.683
PIPAL average SROCC

Average PIPAL SROCC across seven evaluated backbones

What the paper found

FoMo, or Forking Moment, turns a diffusion model’s denoising trajectory into an automatic perceptual-distance label for reference-based image quality assessment. Starting from a reference image, the method injects noise at a selected point in the trajectory and independently denoises afterward: early forks discard shared structure and produce perceptually distant images, while late forks preserve detail and produce close variants. Using FLUX.1-dev with a 50-step schedule, FoMo generated 480k labeled image pairs without human annotation, combining ImageNet references with synthetic references. Instead of conventional MOS regression or 2AFC classification, it trains with a RankNet-style RankBCE objective that compares every pair in a batch, enforcing globally consistent ordering across unrelated reference images. Human validation found a Spearman correlation of 0.970 between fork ordering and perceived similarity, while a cross-reference forced-choice study matched the fork-induced ordering in 90.2% of responses. Across PIPAL, FoMo achieved a 0.733 SROCC using an LPIPS-Alex backbone, compared with 0.577 for KADID-10K supervision, and its average SROCC across seven backbones on PIPAL was 0.683. The approach also worked with Stable Diffusion 1.5, SD-XL, and SD3, although FLUX.1 was strongest. Its main limitation is coverage: because diffusion-generated examples do not naturally represent localized corruption, pixel noise, or purely photometric shifts, performance can weaken on those out-of-distribution distortions.

Original abstract

Reference-based image quality assessment (IQA) metrics aim to reflect how humans perceive the perceptual distance between a pair of images. To learn how the human visual system (HVS) operates, recent reference-based IQA metrics heavily rely on human-annotated data. Mean opinion score (MOS)-based pointwise scoring, which assigns a scalar quality value per image, is preferable for annotation but is prohibitively expensive to collect at scale and is known to be noisy due to inconsistent human judgments. As an alternative, two-alternative forced choice (2AFC) pairwise labels have gained popularity due to their reliability and efficiency, but they capture only relative comparisons between pairs. In this paper, we propose a fully automated data generation pipeline that generates pointwise perceptual distance labels between image pairs without any human annotation. Our approach exploits the generative dynamics of diffusion models as a perceptual distance proxy, where the coarse structure of an image is generated in the early timesteps and the fine details are generated in the later timesteps. Images that fork early in the generation process share only coarse structure and are perceptually far apart; images that fork late differ only in fine detail. We demonstrate that the diffusion trajectory aligns well with the human visual system, and use this forking moment, FoMo, as a reference-grounded distance label to supervise the training of a reference-based IQA metric. The pointwise labels, which support universal comparison between arbitrary image pairs, enable an information-rich training objective. Extensive experiments across diverse backbone architectures confirm the effectiveness of our generation pipeline, outperforming human-annotated datasets in multiple benchmarks.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
01Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
02Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis