NTH

From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models

AuthorsZanyi Wang, Xin Lin, Haodong Li, Dengyang Jiang, Yijiang Li

July 19, 2026 3 min read
Watch on YouTube
The one-line take

This work shows that text-to-image models can do dense vision better by reading out depth, masks, and other per-pixel signals directly from their internal patch grid instead of forcing everything back into RGB.

Key results

33K
Readout head size

Approximate parameters in the shared token-local linear head.

0.063
KITTI depth absRel

ReChannel-9B result on the KITTI monocular-depth benchmark.

5.69
P3M-500-P matting SAD

ReChannel-9B trimap-free matting result.

82.0
RefCOCO referring segmentation cIoU

Average cIoU for ReChannel-9B across the RefCOCO-family splits.

2.48
Speedup over edit pipeline

Times faster than the matched edit-plus-latent-decoding counterpart.

What the paper found

Researchers at UCSD and HKUST propose ReChannel, a way to repurpose Black Forest Labs’ FLUX-Klein text-to-image DiT for dense prediction without generating an image-like target. The method keeps the VAE encoder so RGB inputs remain compatible with the pretrained distribution, adapts the frozen transformer using task-specific LoRA, and applies a shared token-local linear head of about 33K parameters to read each spatial token directly into a pixel-space patch. This removes the target-side VAE decoder and treats the DiT token lattice as a carrier for task-native fields such as depth, surface normals, alpha mattes, referring segmentation, pose heatmaps, and saliency. Across six tasks and more than a dozen benchmarks, ReChannel-9B achieves a KITTI depth absRel of 0.063, trimap-free matting SAD of 5.69 on P3M-500-P, and average referring-segmentation cIoU of 82.0 on the RefCOCO family. It also reaches 79.2 AP on COCO pose estimation. In a matched FLUX-Klein 4B comparison, direct readout runs in 47.7 milliseconds, while an edit-based latent-decoding pipeline takes 118.1 milliseconds, making ReChannel 2.48 times faster while remaining more accurate. Ablations show that LoRA adaptation and a strong pretrained prior are essential, whereas larger spatial heads and full fine-tuning do not improve the compact readout interface.

Original abstract

Large-scale text-to-image models are attractive backbones for dense prediction because RGB generation pretraining learns rich semantic, structural, and geometric priors. Existing generative and editing approaches reuse these priors by casting dense prediction as target generation: annotations such as depth, normals, alpha mattes, masks, and heatmaps are encoded into an RGB-trained VAE latent space and decoded back as image-like targets. We argue this inherits more of the generative output interface than dense prediction requires: unlike RGB synthesis, dense prediction asks for pixel-correct, task-native fields on the same image plane, not new RGB content to be rendered. Our key observation is that a pretrained DiT already organizes RGB inputs through a patch-to-token-to-patch lattice on the image plane, so each token indexes a fixed output patch whose channels can carry task-native quantities instead of RGB appearance. We instantiate this as ReChannel: we keep the VAE encoder for the DiT's input distribution but drop the target-side decoder, adapt the frozen DiT with task LoRA, and map each token to its p x p x K_t pixel-space patch through a shared token-local linear head--about 33K parameters, no spatial mixing. Using FLUX-Klein, we evaluate on six dense prediction tasks and over a dozen benchmarks. This minimal interface sets new state-of-the-art on trimap-free matting, KITTI depth, and referring segmentation, and stays competitive on normals, saliency, and pose. In a matched 4B setting it is more accurate and 2.48x faster than an edit-plus-latent-decode counterpart--dense perception can benefit from generative pretraining without inheriting its output interface.

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis