Hierarchical Denoising For Multi-Step Visual Reasoning
AuthorsZezhong Qian, Xiaowei Chi, Chak-Wing Mak, Tianze Zhou, Ruibin Yuan, Yuhan Rui, Hengzhe Sun, Zhuoqun Wu, Yuming Li, Siyuan Qian, Sirui Han, Shanghang Zhang
Resources
HDR lets video diffusion models plan complex visual actions hierarchically before streaming them efficiently, improving reasoning while sharply reducing inference cost.
Key results
Overall success score on the six-task multi-step visual reasoning benchmark.
Relative improvement in overall success compared with the streaming autoregressive baseline.
Seconds per latent during streaming generation.
Streaming speed advantage relative to bidirectional diffusion.
Full-data success retained when training uses only 2% of the data.
What the paper found
Researchers from Peking University and The Hong Kong University of Science and Technology introduce HDR, or Hierarchical Denoising for Visual Reasoning, built on Wan2.2-5B-TI2V. Instead of forcing streaming autoregressive diffusion to commit frame by frame, HDR organizes video latents into a six-level tree: coarse layers preserve uncertain global plans, while finer layers refine them into concrete visual states. Its SHAP, or Sparse Hierarchical Attention Pattern, restricts each token to local and parent-level contexts, enabling coarse-to-fine planning with sparse KV-cache reuse rather than dense temporal attention. On a 370-video benchmark spanning maze navigation, Tower of Hanoi, one-line drawing, sliding puzzle, Sokoban, and water pouring, HDR raises overall success from 34.22 to 60.29, a 76.2 percent relative gain over CausalForcing, while average progress reaches 89.56 from 76.00. Streaming latency is 0.70 seconds per latent, making HDR 54.2 times faster than bidirectional diffusion during streaming. The model is also data-efficient: with only 2 percent of the training data, it retains 82.9 percent of its full-data success score, versus 52.0 percent for bidirectional diffusion. Robot-maze experiments and the HDR-WAM adaptation to RoboDojo suggest that hierarchical visual planning can transfer to embodied interaction, although remaining failures involve late-stage geometry and state-consistency drift.
Original abstract
Video models are evolving into vision foundation models, yet they still lack human-like multi-step reasoning. Streaming autoregressive diffusion models are efficient but limited in reasoning, while bidirectional diffusion enables global revision with high inference costs due to dense frame-level denoising. Both paradigms struggle to achieve logical consistency and low-latency streaming for complex reasoning tasks. We propose HDR (Hierarchical Denoising for Visual Reasoning), a unified framework that integrates hierarchical latents into causal video generation for multi-step reasoning. HDR organizes video latents into a tree-structured hierarchy, enabling coarse-to-fine reasoning before streaming output. Coarse denoising layers preserve uncertain hypotheses for global planning, while finer layers progressively refine them into concrete visual states. A sparse hierarchical attention pattern (SHAP) further reduces temporal attention costs. We introduce a level-stratified multi-step video reasoning benchmark with out-of-distribution cases, covering six tasks: maze navigation, Tower of Hanoi, one-line drawing, sliding puzzle, Sokoban, and water pouring. Compared with streaming autoregressive diffusion baselines, HDR improves success from 34.22 to 60.29 (76.2% relative gain) and increases average progress from 76.00 to 89.56, demonstrating more consistent reasoning trajectories. HDR maintains low-latency streaming at 0.70 seconds per latent, achieving 54.2 times faster inference than bidirectional diffusion. It also retains 82.9% of full-data performance with only 2% training data, compared with 52.0% for bidirectional diffusion. Real-world robot experiments further demonstrate HDR's potential for physical interaction and world modeling. Project demo: https://hierarchical-diffusion-reasoning.github.io/.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.