Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers
AuthorsMaohua Li, Qirui Li, Yanke Zhou, Yiduo Li, Zhaosheng Chi, Chao Xu, Cuifeng Shen, Yixuan Xu, Hanlin Tang, Kan Liu, Tao Lan, Lin Qu, Shao-Qun Zhang
Resources
This work shows that seemingly uninformative template tokens in diffusion transformers quietly preserve object identity, and that some attention heads can be pruned to save computation with modest quality loss.
Key results
Share of image-to-text attention assigned to structural template tokens in Qwen-Image.
Structural tokens receive 6.4 times the average attention of semantic prompt tokens on GenEval.
Total attention heads ranked for causal transplantation and pruning in Qwen-Image-2512.
Approximate fraction of reverse-ranked heads sufficient to flip object identity in transplantation experiments.
Reduction achieved by training-free late-step pruning.
Point decrease, from 76.1 to 74.7, after the 20% attention-FLOP reduction.
What the paper found
Researchers from Nanjing University, Alibaba Group, and Zhejiang University analyze how chat-formatted text controls diffusion transformers, focusing on Qwen-Image and Qwen-Image-2512. Their causal interpretability framework combines attention decomposition, cross-prompt swaps, head transplantation, and layer-wise masking. The central finding is that trailing structural tokens—such as dialogue delimiters—contain little prompt-specific meaning at the vision-language encoder output, yet become dominant image-to-text attention sinks inside the DiT and function as implicit semantic registers. On GenEval, these template tokens receive 76% of image-to-text attention, and each structural token attracts 6.4 times as much attention as an average prompt token. Causal interventions show that semantics are not transferred directly from prompt tokens into the registers: prompt tokens first inject object identity into image latents, and the structural tokens then read that identity back from the image stream. Transplanting only about 18% of the right attention heads can change an image from an apple to a banana, while heads that read prompt tokens most strongly are largely dispensable. The authors convert this result into a training-free pruning method over 1440 heads: pruning selected heads during late denoising removes 20% of attention FLOPs, reducing GenEval accuracy from 76.1 to 74.7, a 1.4-point drop. The study further separates semantic routing from visual rendering and identifies a depth-wise sequence of early identity commitment, middle-layer propagation, and late refinement.
Original abstract
Text-to-image diffusion transformers (DiTs) jointly process text and image tokens, yet their internal computation during denoising remains poorly understood. We introduce a causal interpretability framework for modern large-scale DiTs that combines attention decomposition with targeted interventions across token spans, heads, and layers. Using it to separate prompt-content tokens from structural template tokens, we find that the structural tokens carry little prompt-specific information at the encoder output. Yet surprisingly, they emerge as dominant image-to-text attention sinks and causally maintain object identity inside the DiT, acting as implicit semantic registers. We show that they acquire this identity indirectly, with prompt semantics first injected into the image latents and then read back into the template tokens rather than transferred directly from the prompt tokens. Inspired by the above findings, we design a training-free pruning rule for DiTs. Heads that attend most strongly to prompt tokens are dispensable, and pruning them removes $20\%$ of attention FLOPs with only a $1.4$-point drop on GenEval. We further reveal how generative computation in DiTs is organized across heads and depth, separating semantic routing from visual synthesis and progressing from identity formation to propagation and refinement. Our work not only reveals that the tokens encoding semantics at input need not be those that maintain it during generation, but also provides a causal view of internal mechanisms in DiTs.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.