NTH

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

AuthorsXiang An, Yin Xie, Feilong Tang, Yunyao Yan, Huajie Tan, Didi Zhu, Changrui Chen, Xiuwei Zhao, Bin Qin, Kaicheng Yang, Yifei Shen, Yuanhan Zhang, Kaichen Zhang, Wenkang Zhang, Zheng Cheng, Nansen Zhang, Chunsheng Wu, Chunjiang Ge, Zimin Ran, Dehua Song, Chunyuan Li, Shikun Feng, Ming Hu, Zhangquan Chen, Junbo Niu, Bo Li, Ziyong Feng, Ziwei Liu, Zongyuan Ge, Jiankang Deng

June 1, 2026 3 min read
Watch on YouTube
The one-line take

LLaVA-OneVision-2 is a next-generation multimodal model that compresses video more intelligently to better understand, ground, and reason over long visual content.

Key results

74.9
JumpScore mAP

LLaVA-OneVision-2-8B achieves 74.9 JumpScore mAP on the new temporal-localization benchmark, versus 30.1 for Qwen3-VL-8B.

44.8
JumpScore gain over Qwen3-VL-8B

On JumpScore, LLaVA-OneVision-2-8B surpasses Qwen3-VL-8B by +44.8 points.

4.3
Video task average gain

Across 18 video tasks, LLaVA-OneVision-2-8B improves the average score by 4.3 points over Qwen3-VL-8B.

5.3
Spatial task average gain

Across 11 spatial tasks, LLaVA-OneVision-2-8B improves the average score by 5.3 points over Qwen3-VL-8B.

15.6
Tracking average J&F gain

Across four tracking benchmarks, LLaVA-OneVision-2-8B improves the average J&F by 15.6 over Qwen3-VL-8B.

8M
Training data scale

The training stack uses approximately 8M re-captioned video samples for pretraining, alongside a 4M-sample spatial corpus.

What the paper found

LLaVA-OneVision-2, from Glint Lab, AIM for Health Lab, and MVP Lab, is a new open vision-language model that pushes the LLaVA-OneVision line toward next-generation perceptual intelligence by replacing frame-centric video input with codec-stream tokenization. Built on the OneVision-Encoder and a Qwen3-8B decoder, it treats compressed video as a continuous bit-cost stream: adaptive GOP boundaries are set from prediction bit-cost, and motion-residual cues select salient 2×2 patch blocks into compact visual canvases, all aligned by shared 3D RoPE and windowed attention. Trained in four stages on roughly 8 million re-captioned video samples plus a 4 million-sample spatial corpus, it preserves image and document capability while improving event-level video reasoning. The model introduces JumpScore, a benchmark for dense repeated motion, where LLaVA-OneVision-2-8B reaches 74.9 mAP, versus 30.1 for Qwen3-VL-8B, a +44.8-point gain. Across 18 video tasks it improves the average score by 4.3 points, across 11 spatial tasks by 5.3 points, and on four tracking benchmarks by 15.6 average J&F. A controlled ablation shows codec-stream inputs outperform uniform frame sampling by 9.7 points on temporal grounding and by 17.3 points on JumpScore, while remaining competitive on long-video QA. The paper’s main claim is that perceptual token allocation should follow coding structure, not elapsed time, and it releases code, data, and model weights as open-source resources.

Original abstract

We introduce LLaVA-OneVision-2 (LLaVA-OV-2), the most capable vision-language model in the LLaVA-OneVision series to date, achieving superior performance across a broad range of multimodal benchmarks. The model builds on a native OneVision-Encoder and incorporates Windowed Attention for efficient local computation while maintaining native resolution. Its key advance is codec-stream tokenization: it treats compressed video as a continuous bit-cost stream, where bit-cost dynamics determine adaptive temporal groups, and motion-residual cues select salient spatial evidence into compact visual canvases. This allocation concentrates a limited token budget on event-bearing content, enabling more stable long-video token compression than fixed groups of pictures. A shared 3D RoPE further places codec canvases, sampled frames, and images in a unified spatiotemporal coordinate system. Furthermore, we build the LLaVA-OV-2 data and training stack around large-scale open supervision: approximately 8M re-captioned video samples for pretraining, a 4M-sample spatial corpus for fine-tuning. We also introduce JumpScore, a temporal-localization benchmark targeting fine-grained grounding in high-frequency, densely repeated motion, a regime underrepresented by existing video evaluations. A standout capability of LLaVA-OV-2 is its unified perception across video understanding, temporal grounding, spatial grounding, and manipulation-trace reasoning. On JumpScore, LLaVA-OneVision-2-8B reaches 74.9 JumpScore mAP, surpassing Qwen3-VL-8B (30.1) by +44.8 points; under matched visual-token budgets on the same benchmark, codec-stream inputs improve temporal grounding over frame sampling by +9.7 points. Across standard benchmarks, LLaVA-OneVision-2-8B further outperforms Qwen3-VL-8B by +4.3 average points on video tasks, +5.3 on spatial tasks, and +15.6 average J&F on tracking tasks.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis