Anchoring Instruction Outside Mask: Exact Reference Caching for Efficient In-Context Diffusion Transformers
AuthorsYangshuai Liu, Zheming Li, Jiaao Li, Kang He, Ziliang Lai, Zhitai Liu, Chengru Song
Resources
This work makes multi-reference diffusion editing much faster by caching visual reference computations without sacrificing instruction following or image fidelity.
Key results
End-to-end acceleration for complete 40-step generation.
Measured speedup when scaling to ten reference images.
Overall score after recovery on OmniContext.
Overall score of the Qwen-Image-Edit-2511 full-attention baseline.
Number of Stage 1 velocity-distillation steps.
Number of Stage 2 student-visited-state updates.
What the paper found
This paper targets a central bottleneck in multi-reference image editing with diffusion transformers: each reference image adds thousands of tokens, and full attention recomputes reference representations at every denoising step. Evaluated on Qwen-Image-Edit-2511, the proposed beyond-mask architecture inserts parameter-free static text anchors during cache construction, allowing reference tokens to access instruction information while keeping their key and value states independent of the evolving target and exactly reusable. The anchors are discarded before sampling, preserving compatibility with optimized kernels such as FlashAttention. Because this architectural change initially harms quality, recovery combines 30k steps of teacher-forced velocity distillation with 500 on-policy updates at states visited by the student. Across OmniContext, GEdit-Bench, and ImgEdit-Bench, the recovered model closely matches full attention: on OmniContext, scored with GPT-4.1, it reaches 8.185 versus 8.119 for the full-attention baseline. For a complete 40-step generation with five reference images, it delivers a 3.92× end-to-end speedup, while scaling to ten references raises the speedup to 5.47×. Static anchors add only 0.3 seconds in the reported latency comparison, showing that instruction-aware exact caching can improve reference fidelity without sacrificing the computational benefits of reuse.
Original abstract
Omnimodal generation is central to a wide range of content creation and editing applications. In-context conditioning is essential to this paradigm. It allows diffusion transformers to process text instructions and visual references in a shared attention sequence. However, each reference image introduces thousands of tokens. Computation therefore grows rapidly with the number of references. Existing methods reduce computation through structured sparse attention, which limits interactions between reference and target tokens. This structure also makes the reference K and V independent of the denoising target, allowing them to be computed once and reused across steps. However, it blocks visual references from attending to the text instruction. This substantially degrades instruction following and reference fidelity in multi-reference editing. To resolve this conflict, we jointly redesign the token sequence and attention mask. Our beyond-mask design uses static text anchors to connect the instruction to the reference branch. It preserves exact K and V reuse without adding parameters. However, this direct architectural conversion degrades generation quality. We recover the lost performance through teacher-forced velocity distillation, followed by a short on-policy stage in which the teacher supervises student-visited states. To our knowledge, this is the first use of on-policy distillation for architectural recovery in diffusion models. Across three image-editing benchmarks, our method matches full-attention generation quality. With five reference images, it accelerates the complete 40-step denoising process by 3.92x, while static text anchors introduce negligible runtime overhead; the speedup reaches 5.47x at ten references in our scaling study.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.