NTH

Let RGB Be the Language of Vision

AuthorsTiming Yang, Jinrui Yang, Xinlong Li, Yuhan Wang, Haoran Li, Yanqing Liu, Guoyizhe Wei, Jixuan Ying, Chen Wei, Rama Chellappa, Yuyin Zhou, Cihang Xie, Alan Yuille, Feng Wang

July 27, 2026 2 min read
Watch on YouTube
The one-line take

RINO treats every visual signal as an image, allowing one shared model to understand and generate across tasks like segmentation, depth estimation, and pose-guided synthesis.

Key results

25
RINO evaluation scope

Number of vision tasks spanning estimation, 3D geometry, grounding, and conditioned generation.

20B
Qwen-Image-Edit model size

Parameter count of the main open image-editing backbone used by RINO.

0.938
DIODE-indoor depth δ1

Zero-shot Qwen-Image-Edit score, compared with Depth Anything V2 at 0.952.

28.0
LongCat instance-map AP

COCO2017 instance-map-conditioned generation score, above InstanceDiffusion at 27.1.

14.90
Qwen Canny F1

Zero-shot edge controllability score on MultiGen-20M, versus ControlNet++ at 37.04.

What the paper found

Researchers from Johns Hopkins University, UC Santa Cruz, Carnegie Mellon University, and Rice University, including Alan Yuille and Rama Chellappa, introduce RINO, or RGB In and RGB Out, a unified vision interface that converts masks, depth, surface normals, poses, layouts, and edges into RGB images. A frozen image editor then handles both perception and generation as RGB-to-RGB editing, using only parameter-free color conversion and no task-specific heads, encoders, adapters, or fine-tuning. Evaluated across 25 tasks, RINO-Zero runs on open models including Qwen-Image-Edit, a 20B-parameter editor, LongCat-Image-Edit from Meituan, and FireRed-Image-Edit. On DIODE-indoor depth estimation, Qwen reaches a δ1 score of 0.938, approaching the specialist Depth Anything V2 at 0.952. For instance-map-conditioned generation on COCO2017, LongCat achieves an AP of 28.0, exceeding the task-trained InstanceDiffusion reference at 27.1 while producing better FID. The approach also supports semantic segmentation, panoptic segmentation, referring expression grounding, pose estimation, and depth-, segmentation-, pose-, instance-, and Canny-conditioned synthesis. Its limitations expose where specialization still matters: on Canny control, Qwen obtains an F1 of 14.90 versus ControlNet++ at 37.04, and recognition and instance association remain weaker than spatial mask quality. Overall, the paper argues that RGB can function as a shared visual language, analogous to text in language models, while identifying instruction-tuning on RGB-formatted structured signals as the next step.

Original abstract

This work introduces a unified formulation for vision models, where diverse forms of visual information beyond natural images, such as masks, depth maps, and other structured visual signals, are all represented as RGB images, while general visual tasks can be converted into a common RGB-to-RGB image editing problem. In this paradigm, different types of visual information internally share the same encoding and decoding architecture and parameters as natural images, enabling a single model to transfer across tasks through a unified visual interface, in a way analogous to how language models operate over text. We refer to this formulation as RGB In and RGB Out (RINO). Built upon a generic image editing backbone without task-specific fine-tuning, RINO demonstrates robust and competitive zero-shot performance on both dense understanding tasks such as segmentation and depth estimation (where we unify outputs as RGB), and dense-conditioned generation tasks such as pose-to-image generation (where we unify inputs as RGB). We hope this study provides useful insights toward general unified vision-language systems, where diverse visual tasks can be expressed, interpreted, and solved through a shared visual language. Code is available at https://github.com/yangtiming/RINO.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis