NTH

Anchoring Instruction Outside Mask: Exact Reference Caching for Efficient In-Context Diffusion Transformers

AuthorsYangshuai Liu, Zheming Li, Jiaao Li, Kang He, Ziliang Lai, Zhitai Liu, Chengru Song

September 2, 2026 2 min read
Watch on YouTube
The one-line take

This work makes multi-reference diffusion editing much faster by caching visual reference computations without sacrificing instruction following or image fidelity.

Key results

3.92×
Five-reference speedup

End-to-end acceleration for complete 40-step generation.

5.47×
Ten-reference scaling speedup

Measured speedup when scaling to ten reference images.

8.185
Text-anchor OmniContext score

Overall score after recovery on OmniContext.

8.119
Full-attention OmniContext score

Overall score of the Qwen-Image-Edit-2511 full-attention baseline.

30k
Teacher-forced recovery steps

Number of Stage 1 velocity-distillation steps.

500
On-policy recovery updates

Number of Stage 2 student-visited-state updates.

What the paper found

This paper targets a central bottleneck in multi-reference image editing with diffusion transformers: each reference image adds thousands of tokens, and full attention recomputes reference representations at every denoising step. Evaluated on Qwen-Image-Edit-2511, the proposed beyond-mask architecture inserts parameter-free static text anchors during cache construction, allowing reference tokens to access instruction information while keeping their key and value states independent of the evolving target and exactly reusable. The anchors are discarded before sampling, preserving compatibility with optimized kernels such as FlashAttention. Because this architectural change initially harms quality, recovery combines 30k steps of teacher-forced velocity distillation with 500 on-policy updates at states visited by the student. Across OmniContext, GEdit-Bench, and ImgEdit-Bench, the recovered model closely matches full attention: on OmniContext, scored with GPT-4.1, it reaches 8.185 versus 8.119 for the full-attention baseline. For a complete 40-step generation with five reference images, it delivers a 3.92× end-to-end speedup, while scaling to ten references raises the speedup to 5.47×. Static anchors add only 0.3 seconds in the reported latency comparison, showing that instruction-aware exact caching can improve reference fidelity without sacrificing the computational benefits of reuse.

Original abstract

Omnimodal generation is central to a wide range of content creation and editing applications. In-context conditioning is essential to this paradigm. It allows diffusion transformers to process text instructions and visual references in a shared attention sequence. However, each reference image introduces thousands of tokens. Computation therefore grows rapidly with the number of references. Existing methods reduce computation through structured sparse attention, which limits interactions between reference and target tokens. This structure also makes the reference K and V independent of the denoising target, allowing them to be computed once and reused across steps. However, it blocks visual references from attending to the text instruction. This substantially degrades instruction following and reference fidelity in multi-reference editing. To resolve this conflict, we jointly redesign the token sequence and attention mask. Our beyond-mask design uses static text anchors to connect the instruction to the reference branch. It preserves exact K and V reuse without adding parameters. However, this direct architectural conversion degrades generation quality. We recover the lost performance through teacher-forced velocity distillation, followed by a short on-policy stage in which the teacher supervises student-visited states. To our knowledge, this is the first use of on-policy distillation for architectural recovery in diffusion models. Across three image-editing benchmarks, our method matches full-attention generation quality. With five reference images, it accelerates the complete 40-step denoising process by 3.92x, while static text anchors introduce negligible runtime overhead; the speedup reaches 5.47x at ten references in our scaling study.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis