NTH

From Pixels to Words -- Towards Native One-Vision Models at Scale

AuthorsHaiwen Diao, Jiahao Wang, Penghao Wu, Yuhao Dong, Yuwei Niu, Yue Zhu, Zhongang Cai, Weichen Fan, Linjun Dai, Silei Wu, Xuanyu Zheng, Mingxuan Li, Yuanhan Zhang, Bo Li, Hanming Deng, Huchuan Lu, Quan Wang, Lei Yang, Lewei Lu, Dahua Lin, Ziwei Liu

June 26, 2026 2 min read
Watch on YouTube
The one-line take

This paper introduces a native vision-language model that learns pixel-to-word connections end-to-end without separate encoders or adapters, aiming to make multimodal systems more unified and competitive at scale.

Key results

20M
pretraining image-text pairs

large-scale image-text pairs used in stage 1

60M
mid-training multimodal samples

multimodal samples used in stage 2

6M
instruction-tuning samples

high-quality image/video instruction data used in stage 3

36K
context length

maximum context length reached during training

4096^2
visual resolution

maximum image resolution used during training

128
video frames

maximum sampled frames per video during training

What the paper found

NEO-ov, from SenseTime Research and NTU’s S-Lab, is a fully native one-vision foundation model that removes the conventional vision encoder, projector, and post-hoc fusion stack used by modular VLMs such as Qwen3-VL and InternVL3.5. Instead, it serializes images, multi-image sets, and video frames directly into one decoder-only backbone with native patch embeddings, T-H-W decoupled attention, and Native RoPE, so pixel-word and pixel-pixel interactions begin at the earliest layers. The model is trained in three stages on 20M image-text pairs, nearly 60M multimodal samples, and 6M high-quality instruction samples, with context length expanded from 16K to 36K and visual resolution up to 4096^2; videos are sampled up to 128 frames. At 2B scale, NEO-ov reaches 54.7 on MMMU, 80.0 on MMBench, 91.2 on DocVQA, and 81.2 on OCRBench, exceeding prior native models and approaching modular counterparts. At 8B scale, it scores 68.1 on MMMU, 85.1 on MMBench, 91.9 on DocVQA, 81.6 on OCRBench, 67.4 on VideoMME, 70.7 on MVBench, and 90.0 on Mindcube, showing particularly strong gains in spatial intelligence, where it attains 64.8 on VSI-Bench and 90.0 on Mindcube. The main takeaway is that a monolithic, encoder-free architecture can scale to competitive multimodal reasoning while improving fine-grained spatial perception and cross-frame correspondence.

Original abstract

Current vision-language models (VLMs) typically stitch together separate image encoders and language decoders via multi-stage alignment, a modular framework that inevitably fragments pixel-level signals across frames and scatters early pixel-word interactions. In parallel, native VLMs, despite impressive performance on single images, remain largely unexplored in multi-image, video understanding, and spatial intelligence. Hence, we introduce NEO-ov, a native foundation model that learns cross-frame and pixel-word correspondence end-to-end, without any external encoders, auxiliary adapters, or post-hoc fusion. By eliminating module boundaries entirely, NEO-ov enables fine-grained and unified spatiotemporal modeling to emerge natively inside the model. Notably, NEO-ov largely narrows the gap to modular counterparts while excelling at fine-grained visual perception, validating that native "one-vision" architectures are not only feasible but competitive at scale. Beyond empirical performance, we unveil systematic architectural analyses and detailed training recipes to facilitate subsequent native multimodal modeling. Our code and models are publicly available at: https://github.com/EvolvingLMMs-Lab/NEO.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis