NTH

Multi-Turn Reflective Masking Elicits Reasoning in Mask Diffusion Models

AuthorsYanming Zhang, Yihan Bian, Jingyuan Qi, Yuguang Yao, Lifu Huang, Tianyi Zhou

June 24, 2026 2 min read
Watch on YouTube
The one-line take

This paper teaches mask diffusion models to “think again” by revisiting and locally revising their own outputs across multiple turns, enabling stronger reasoning and refinement without starting over.

Key results

6
trajectory length

training and inference horizon used in the proposed RM procedure

5
training time

hours to train the three-task setup on two H100 GPUs

2
GPU count

NVIDIA H100 80GB GPUs used for training

0.81M
Sudoku model size

parameter count of the lightweight Sudoku MDM

24.8
MATH500 accuracy

RM performance on MATH500 versus 22.4 for Vanilla SFT

What the paper found

Multi Turn Reflective Masking Elicits Reasoning in Mask Diffusion Models, from authors at the University of Maryland, Virginia Tech, Intuit, UC Davis, and MBZUAI, argues that mask diffusion models can reason by revisiting and revising prior tokens in place rather than regenerating sequences autoregressively. The paper introduces Reflective Masking, a lightweight post-training method that turns each position into a keep, re-mask, or reveal decision driven by the model’s own per-position probabilities, and adds History Reference, a parameter-free rotary accumulation of intermediate denoising states that stabilizes multi-turn revision. With a trajectory length of 6 and no architectural changes, the method trains in about 5 hours on 2 NVIDIA H100 80GB GPUs. On ImgEdit with Lumina-DiMOO, RM raises Edit Precision from 71.84 to 99.73, SSIM from 0.6570 to 0.9744, and VQAScore from 81.61 to 85.17, while improving user preference to 68.2. On Sudoku, a 0.81M-parameter four-layer Transformer, full RM plus History Reference reaches 93.4% exact accuracy and 93.6% valid rate, cutting replay mistakes to 0.03 and conflict cells to 0.236 per board. On text reasoning with LLaDA, RM improves MATH500 from 22.4% to 24.8%, MBPP from 30.6% to 39.4%, and ARC-Challenge from 81.3% to 86.1%, showing that iterative in-place masking can function as a distinct test-time scaling mechanism for diffusion-based generation.

Original abstract

While reasoning on autoregressive (AR) models is often performed by chain-of-thought reasoning and reflection, their refinement of previous outputs still relies on fully sequential generation, even when only local edits are needed. In contrast, the masking mechanism in Mask Diffusion Models (MDMs) naturally supports explicit local edits on previous outputs, allowing selective refinement without discarding previous answers and generating another from scratch. While this property more closely aligns with how humans correct mistakes by iterative local refinement, existing MDMs do not support multi-turn masking and denoising. We propose Reflective Masking (RM), which elicits such an intrinsic reasoning capability in MDMs via lightweight post-training. RM provides a native test-time scaling, where an MDM iteratively revisits and revises its prior outputs based on evolving context. To exploit insights from previous turns like AR reasoning, we further introduce History Reference, a parameter-free mechanism that leverages intermediate denoising states during revision. Our approach requires no architectural changes and is easily applicable to existing MDMs. Across diverse tasks and modalities, including text generation, Sudoku, and image editing, Reflective Masking consistently outperforms standard masking-based baselines and demonstrates strong generality, positioning RM as a fundamental primitive for reasoning on MDMs.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis