NTH

Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases Them

AuthorsWoojung Han, Seil Kang, Youngjun Jun, Min-Hung Chen, Fu-En Yang, Seong Jae Hwang

June 9, 2026 2 min read
Watch on YouTube
The one-line take

This paper finds that early diffusion steps can preserve better physics than later ones, and introduces a lightweight method to lock in that motion prior while still producing high-quality video.

Key results

18%
phase coherence drop

Low-frequency phase coherence decreases from step 2 to step 50 during denoising

6.2
Physics-IQ average gain

Average Physics-IQ improvement from PhaseLock across backbones

36.0
CogVideoX-5B score

Physics-IQ score after applying PhaseLock to CogVideoX-5B

28.7
Wan 2.1 score

Physics-IQ score after applying PhaseLock to Wan 2.1

32.0
LTX-Video score

Physics-IQ score after applying PhaseLock to LTX-Video

1.06x
Inference time overhead

PhaseLock runtime overhead relative to the baseline

What the paper found

This paper from Yonsei University and NVIDIA introduces PhaseLock, a training-free image-to-video framework that argues physical motion is already present after just 2 denoising steps and is then progressively erased by visual refinement. Using CogVideoX-5B and Wan 2.1, the authors show that a 2-step output can score better on physics than a 50-step output, while spectral analysis reveals an approximately 18% drop in low-frequency phase coherence from step 2 to step 50 even though magnitude stays nearly stable. PhaseLock extracts a motion prior from the 2-step latent and re-injects it during full 50-step generation through Latent Delta Guidance, a spatial-domain constraint on inter-frame latent differences that avoids explicit FFT surgery. On the Physics-IQ benchmark, PhaseLock improves base models by an average of 6.2 points, raising CogVideoX-5B from 30.8 to 36.0, LTX-Video from 26.4 to 32.0, and Wan 2.1 from 20.9 to 28.7, while preserving visual quality on VBench and adding only 1.06× time and 1.02× memory overhead. The method also outperforms direct frequency manipulation baselines, including low-frequency phase injection and full phase substitution, which collapse to scores near 14, and it matches WMReward’s physics alignment with far less cost than WMReward’s roughly 5× inference-time overhead. Overall, the paper’s key claim is that physics-aware generation should preserve early motion priors rather than prolong denoising, because the extra steps improve photorealism while systematically eroding phase-coded dynamics.

Original abstract

Image-to-Video diffusion models leverage input images to generate visually stunning content, yet frequently produce motion that violates physical laws. We reveal a surprising finding: a 2-step generation often exhibits better physical consistency than a 50-step output from the same model. Through spectral analysis, we trace this to phase erosion during denoising; the phase degrades significantly (dropping by $\approx 18\%$ from step 2 to step 50), whereas the magnitude remains relatively stable. Building on this insight, we propose PhaseLock, a training-free framework that preserves the valid motion priors from few-step inference throughout the denoising trajectory. Rather than relying on full-step inference for physical consistency, PhaseLock extracts a motion prior from just 2 steps and enforces it onto high-fidelity generation via Latent Delta Guidance. Our approach effectively mitigates phase degradation, improving physical consistency by an average of 6.2 points across diverse models while largely maintaining visual fidelity, with negligible overhead ($1.06\times$ time, $1.02\times$ memory) and reduced reliance on expensive external guidance methods ($\sim5\times$ time).

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis