NTH

Appearance Pointers -- Multimodal Region Control of Diffusion Transformers

AuthorsRahul Sajnani, Yulia Gryaditskaya, Radomír Měch, Srinath Sridhar, Matheus Gadelha

July 26, 2026 2 min read
Watch on YouTube
The one-line take

A new diffusion-transformer interface lets users precisely control what appears in different image regions using text or image guidance.

Key results

37K
AppearancePointers-37K dataset

Synthetic dataset used for training and evaluation with regional text, image, and mask conditioning.

500
Evaluation benchmark size

Images evaluated under text-only and image-only regional conditioning.

56.09
Text-region DINO-I

Semantic region-fidelity score on text-conditioned generation.

69.31
Image-region DINO-I

Identity-preservation score on image-conditioned generation.

40.97
Image-region MIoU

Region-mask adherence score for image-conditioned generation.

400M
Added module parameters

Approximate size of the Region Correspondence and Region Aggregation modules, representing a 3.33% increase over the base model.

What the paper found

Researchers from Brown University and Adobe Research introduce Appearance Pointers, a modality-agnostic control interface for Diffusion Transformers, demonstrated with Black Forest Labs’ FLUX model. The method aligns text descriptions, reference images, and user-defined region masks through a Region Correspondence Transformer, then compresses multiple regional signals into spatially aggregated pointer tokens. These pointers direct the DiT toward the correct appearance cues and locations during a single denoising pass, while region-contour guidance improves boundary precision. The same model supports fine and sparse layout generation, multi-subject insertion, pose control, editing, and simultaneous image-plus-text conditioning without retraining the base model from scratch. The authors create the synthetic AppearancePointers-37K dataset and evaluate on a 500-image benchmark averaging five regions per image. For text-conditioned regions, the method reaches DINO-I 56.09 and CLIP-IQA 95.02; for image-conditioned regions, it achieves DINO-I 69.31, CLIP-I 93.29, and MIoU 40.97, surpassing MSDiffusion and DreamRenderer on identity and spatial adherence. The added correspondence and aggregation modules contain approximately 400M parameters, a 3.33% increase over the base model, while pointer computation occurs once per generation rather than at every diffusion step. Ablations show that removing region aggregation reduces DINO-I to 54.47, and removing contour guidance reduces MIoU to 35.81. Limitations include weaker fine-grained face identity preservation and degradation when conditioning on about 10 or more regions.

Original abstract

Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements that cannot be reliably achieved through text prompting alone. Diffusion Transformers (DiTs) can natively ingest heterogeneous tokens stemming from texts and images, but they lack mechanisms for determining where and how these tokens should influence the output. We introduce appearance pointers, compact tokens that guide DiTs toward the correct appearance cues at the correct spatial locations by aligning text or image inputs with user-specified masks. Appearance pointers are produced by a region correspondence network and refined through a spatial aggregation mechanism, enabling the model to handle multiple regional descriptions without significantly increasing token load. Our approach introduces the first modality-agnostic interface for localized multimodal control in a DiT without retraining the base model from scratch. Across a range of metrics, our single model reaches or surpasses the performance of modality-specific state of the art methods, offering a simple and extensible path toward precise, region-aware, multimodal guidance in generative image synthesis.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis