From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models
AuthorsZanyi Wang, Xin Lin, Haodong Li, Dengyang Jiang, Yijiang Li
Resources
This work shows that text-to-image models can do dense vision better by reading out depth, masks, and other per-pixel signals directly from their internal patch grid instead of forcing everything back into RGB.
Key results
Approximate parameters in the shared token-local linear head.
ReChannel-9B result on the KITTI monocular-depth benchmark.
ReChannel-9B trimap-free matting result.
Average cIoU for ReChannel-9B across the RefCOCO-family splits.
Times faster than the matched edit-plus-latent-decoding counterpart.
What the paper found
Researchers at UCSD and HKUST propose ReChannel, a way to repurpose Black Forest Labs’ FLUX-Klein text-to-image DiT for dense prediction without generating an image-like target. The method keeps the VAE encoder so RGB inputs remain compatible with the pretrained distribution, adapts the frozen transformer using task-specific LoRA, and applies a shared token-local linear head of about 33K parameters to read each spatial token directly into a pixel-space patch. This removes the target-side VAE decoder and treats the DiT token lattice as a carrier for task-native fields such as depth, surface normals, alpha mattes, referring segmentation, pose heatmaps, and saliency. Across six tasks and more than a dozen benchmarks, ReChannel-9B achieves a KITTI depth absRel of 0.063, trimap-free matting SAD of 5.69 on P3M-500-P, and average referring-segmentation cIoU of 82.0 on the RefCOCO family. It also reaches 79.2 AP on COCO pose estimation. In a matched FLUX-Klein 4B comparison, direct readout runs in 47.7 milliseconds, while an edit-based latent-decoding pipeline takes 118.1 milliseconds, making ReChannel 2.48 times faster while remaining more accurate. Ablations show that LoRA adaptation and a strong pretrained prior are essential, whereas larger spatial heads and full fine-tuning do not improve the compact readout interface.
Original abstract
Large-scale text-to-image models are attractive backbones for dense prediction because RGB generation pretraining learns rich semantic, structural, and geometric priors. Existing generative and editing approaches reuse these priors by casting dense prediction as target generation: annotations such as depth, normals, alpha mattes, masks, and heatmaps are encoded into an RGB-trained VAE latent space and decoded back as image-like targets. We argue this inherits more of the generative output interface than dense prediction requires: unlike RGB synthesis, dense prediction asks for pixel-correct, task-native fields on the same image plane, not new RGB content to be rendered. Our key observation is that a pretrained DiT already organizes RGB inputs through a patch-to-token-to-patch lattice on the image plane, so each token indexes a fixed output patch whose channels can carry task-native quantities instead of RGB appearance. We instantiate this as ReChannel: we keep the VAE encoder for the DiT's input distribution but drop the target-side decoder, adapt the frozen DiT with task LoRA, and map each token to its p x p x K_t pixel-space patch through a shared token-local linear head--about 33K parameters, no spatial mixing. Using FLUX-Klein, we evaluate on six dense prediction tasks and over a dozen benchmarks. This minimal interface sets new state-of-the-art on trimap-free matting, KITTI depth, and referring segmentation, and stays competitive on normals, saliency, and pose. In a matched 4B setting it is more accurate and 2.48x faster than an edit-plus-latent-decode counterpart--dense perception can benefit from generative pretraining without inheriting its output interface.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.