Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases Them
AuthorsWoojung Han, Seil Kang, Youngjun Jun, Min-Hung Chen, Fu-En Yang, Seong Jae Hwang
Resources
This paper finds that early diffusion steps can preserve better physics than later ones, and introduces a lightweight method to lock in that motion prior while still producing high-quality video.
Key results
Low-frequency phase coherence decreases from step 2 to step 50 during denoising
Average Physics-IQ improvement from PhaseLock across backbones
Physics-IQ score after applying PhaseLock to CogVideoX-5B
Physics-IQ score after applying PhaseLock to Wan 2.1
Physics-IQ score after applying PhaseLock to LTX-Video
PhaseLock runtime overhead relative to the baseline
What the paper found
This paper from Yonsei University and NVIDIA introduces PhaseLock, a training-free image-to-video framework that argues physical motion is already present after just 2 denoising steps and is then progressively erased by visual refinement. Using CogVideoX-5B and Wan 2.1, the authors show that a 2-step output can score better on physics than a 50-step output, while spectral analysis reveals an approximately 18% drop in low-frequency phase coherence from step 2 to step 50 even though magnitude stays nearly stable. PhaseLock extracts a motion prior from the 2-step latent and re-injects it during full 50-step generation through Latent Delta Guidance, a spatial-domain constraint on inter-frame latent differences that avoids explicit FFT surgery. On the Physics-IQ benchmark, PhaseLock improves base models by an average of 6.2 points, raising CogVideoX-5B from 30.8 to 36.0, LTX-Video from 26.4 to 32.0, and Wan 2.1 from 20.9 to 28.7, while preserving visual quality on VBench and adding only 1.06× time and 1.02× memory overhead. The method also outperforms direct frequency manipulation baselines, including low-frequency phase injection and full phase substitution, which collapse to scores near 14, and it matches WMReward’s physics alignment with far less cost than WMReward’s roughly 5× inference-time overhead. Overall, the paper’s key claim is that physics-aware generation should preserve early motion priors rather than prolong denoising, because the extra steps improve photorealism while systematically eroding phase-coded dynamics.
Original abstract
Image-to-Video diffusion models leverage input images to generate visually stunning content, yet frequently produce motion that violates physical laws. We reveal a surprising finding: a 2-step generation often exhibits better physical consistency than a 50-step output from the same model. Through spectral analysis, we trace this to phase erosion during denoising; the phase degrades significantly (dropping by $\approx 18\%$ from step 2 to step 50), whereas the magnitude remains relatively stable. Building on this insight, we propose PhaseLock, a training-free framework that preserves the valid motion priors from few-step inference throughout the denoising trajectory. Rather than relying on full-step inference for physical consistency, PhaseLock extracts a motion prior from just 2 steps and enforces it onto high-fidelity generation via Latent Delta Guidance. Our approach effectively mitigates phase degradation, improving physical consistency by an average of 6.2 points across diverse models while largely maintaining visual fidelity, with negligible overhead ($1.06\times$ time, $1.02\times$ memory) and reduced reliance on expensive external guidance methods ($\sim5\times$ time).
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.