Let RGB Be the Language of Vision
AuthorsTiming Yang, Jinrui Yang, Xinlong Li, Yuhan Wang, Haoran Li, Yanqing Liu, Guoyizhe Wei, Jixuan Ying, Chen Wei, Rama Chellappa, Yuyin Zhou, Cihang Xie, Alan Yuille, Feng Wang
RINO treats every visual signal as an image, allowing one shared model to understand and generate across tasks like segmentation, depth estimation, and pose-guided synthesis.
Key results
Number of vision tasks spanning estimation, 3D geometry, grounding, and conditioned generation.
Parameter count of the main open image-editing backbone used by RINO.
Zero-shot Qwen-Image-Edit score, compared with Depth Anything V2 at 0.952.
COCO2017 instance-map-conditioned generation score, above InstanceDiffusion at 27.1.
Zero-shot edge controllability score on MultiGen-20M, versus ControlNet++ at 37.04.
What the paper found
Researchers from Johns Hopkins University, UC Santa Cruz, Carnegie Mellon University, and Rice University, including Alan Yuille and Rama Chellappa, introduce RINO, or RGB In and RGB Out, a unified vision interface that converts masks, depth, surface normals, poses, layouts, and edges into RGB images. A frozen image editor then handles both perception and generation as RGB-to-RGB editing, using only parameter-free color conversion and no task-specific heads, encoders, adapters, or fine-tuning. Evaluated across 25 tasks, RINO-Zero runs on open models including Qwen-Image-Edit, a 20B-parameter editor, LongCat-Image-Edit from Meituan, and FireRed-Image-Edit. On DIODE-indoor depth estimation, Qwen reaches a δ1 score of 0.938, approaching the specialist Depth Anything V2 at 0.952. For instance-map-conditioned generation on COCO2017, LongCat achieves an AP of 28.0, exceeding the task-trained InstanceDiffusion reference at 27.1 while producing better FID. The approach also supports semantic segmentation, panoptic segmentation, referring expression grounding, pose estimation, and depth-, segmentation-, pose-, instance-, and Canny-conditioned synthesis. Its limitations expose where specialization still matters: on Canny control, Qwen obtains an F1 of 14.90 versus ControlNet++ at 37.04, and recognition and instance association remain weaker than spatial mask quality. Overall, the paper argues that RGB can function as a shared visual language, analogous to text in language models, while identifying instruction-tuning on RGB-formatted structured signals as the next step.
Original abstract
This work introduces a unified formulation for vision models, where diverse forms of visual information beyond natural images, such as masks, depth maps, and other structured visual signals, are all represented as RGB images, while general visual tasks can be converted into a common RGB-to-RGB image editing problem. In this paradigm, different types of visual information internally share the same encoding and decoding architecture and parameters as natural images, enabling a single model to transfer across tasks through a unified visual interface, in a way analogous to how language models operate over text. We refer to this formulation as RGB In and RGB Out (RINO). Built upon a generic image editing backbone without task-specific fine-tuning, RINO demonstrates robust and competitive zero-shot performance on both dense understanding tasks such as segmentation and depth estimation (where we unify outputs as RGB), and dense-conditioned generation tasks such as pose-to-image generation (where we unify inputs as RGB). We hope this study provides useful insights toward general unified vision-language systems, where diverse visual tasks can be expressed, interpreted, and solved through a shared visual language. Code is available at https://github.com/yangtiming/RINO.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.