Multi-Turn Reflective Masking Elicits Reasoning in Mask Diffusion Models
AuthorsYanming Zhang, Yihan Bian, Jingyuan Qi, Yuguang Yao, Lifu Huang, Tianyi Zhou
Resources
This paper teaches mask diffusion models to “think again” by revisiting and locally revising their own outputs across multiple turns, enabling stronger reasoning and refinement without starting over.
Key results
training and inference horizon used in the proposed RM procedure
hours to train the three-task setup on two H100 GPUs
NVIDIA H100 80GB GPUs used for training
parameter count of the lightweight Sudoku MDM
RM performance on MATH500 versus 22.4 for Vanilla SFT
What the paper found
Multi Turn Reflective Masking Elicits Reasoning in Mask Diffusion Models, from authors at the University of Maryland, Virginia Tech, Intuit, UC Davis, and MBZUAI, argues that mask diffusion models can reason by revisiting and revising prior tokens in place rather than regenerating sequences autoregressively. The paper introduces Reflective Masking, a lightweight post-training method that turns each position into a keep, re-mask, or reveal decision driven by the model’s own per-position probabilities, and adds History Reference, a parameter-free rotary accumulation of intermediate denoising states that stabilizes multi-turn revision. With a trajectory length of 6 and no architectural changes, the method trains in about 5 hours on 2 NVIDIA H100 80GB GPUs. On ImgEdit with Lumina-DiMOO, RM raises Edit Precision from 71.84 to 99.73, SSIM from 0.6570 to 0.9744, and VQAScore from 81.61 to 85.17, while improving user preference to 68.2. On Sudoku, a 0.81M-parameter four-layer Transformer, full RM plus History Reference reaches 93.4% exact accuracy and 93.6% valid rate, cutting replay mistakes to 0.03 and conflict cells to 0.236 per board. On text reasoning with LLaDA, RM improves MATH500 from 22.4% to 24.8%, MBPP from 30.6% to 39.4%, and ARC-Challenge from 81.3% to 86.1%, showing that iterative in-place masking can function as a distinct test-time scaling mechanism for diffusion-based generation.
Original abstract
While reasoning on autoregressive (AR) models is often performed by chain-of-thought reasoning and reflection, their refinement of previous outputs still relies on fully sequential generation, even when only local edits are needed. In contrast, the masking mechanism in Mask Diffusion Models (MDMs) naturally supports explicit local edits on previous outputs, allowing selective refinement without discarding previous answers and generating another from scratch. While this property more closely aligns with how humans correct mistakes by iterative local refinement, existing MDMs do not support multi-turn masking and denoising. We propose Reflective Masking (RM), which elicits such an intrinsic reasoning capability in MDMs via lightweight post-training. RM provides a native test-time scaling, where an MDM iteratively revisits and revises its prior outputs based on evolving context. To exploit insights from previous turns like AR reasoning, we further introduce History Reference, a parameter-free mechanism that leverages intermediate denoising states during revision. Our approach requires no architectural changes and is easily applicable to existing MDMs. Across diverse tasks and modalities, including text generation, Sudoku, and image editing, Reflective Masking consistently outperforms standard masking-based baselines and demonstrates strong generality, positioning RM as a fundamental primitive for reasoning on MDMs.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.