LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence
AuthorsXiang An, Yin Xie, Feilong Tang, Yunyao Yan, Huajie Tan, Didi Zhu, Changrui Chen, Xiuwei Zhao, Bin Qin, Kaicheng Yang, Yifei Shen, Yuanhan Zhang, Kaichen Zhang, Wenkang Zhang, Zheng Cheng, Nansen Zhang, Chunsheng Wu, Chunjiang Ge, Zimin Ran, Dehua Song, Chunyuan Li, Shikun Feng, Ming Hu, Zhangquan Chen, Junbo Niu, Bo Li, Ziyong Feng, Ziwei Liu, Zongyuan Ge, Jiankang Deng
Resources
LLaVA-OneVision-2 is a next-generation multimodal model that compresses video more intelligently to better understand, ground, and reason over long visual content.
Key results
LLaVA-OneVision-2-8B achieves 74.9 JumpScore mAP on the new temporal-localization benchmark, versus 30.1 for Qwen3-VL-8B.
On JumpScore, LLaVA-OneVision-2-8B surpasses Qwen3-VL-8B by +44.8 points.
Across 18 video tasks, LLaVA-OneVision-2-8B improves the average score by 4.3 points over Qwen3-VL-8B.
Across 11 spatial tasks, LLaVA-OneVision-2-8B improves the average score by 5.3 points over Qwen3-VL-8B.
Across four tracking benchmarks, LLaVA-OneVision-2-8B improves the average J&F by 15.6 over Qwen3-VL-8B.
The training stack uses approximately 8M re-captioned video samples for pretraining, alongside a 4M-sample spatial corpus.
What the paper found
LLaVA-OneVision-2, from Glint Lab, AIM for Health Lab, and MVP Lab, is a new open vision-language model that pushes the LLaVA-OneVision line toward next-generation perceptual intelligence by replacing frame-centric video input with codec-stream tokenization. Built on the OneVision-Encoder and a Qwen3-8B decoder, it treats compressed video as a continuous bit-cost stream: adaptive GOP boundaries are set from prediction bit-cost, and motion-residual cues select salient 2×2 patch blocks into compact visual canvases, all aligned by shared 3D RoPE and windowed attention. Trained in four stages on roughly 8 million re-captioned video samples plus a 4 million-sample spatial corpus, it preserves image and document capability while improving event-level video reasoning. The model introduces JumpScore, a benchmark for dense repeated motion, where LLaVA-OneVision-2-8B reaches 74.9 mAP, versus 30.1 for Qwen3-VL-8B, a +44.8-point gain. Across 18 video tasks it improves the average score by 4.3 points, across 11 spatial tasks by 5.3 points, and on four tracking benchmarks by 15.6 average J&F. A controlled ablation shows codec-stream inputs outperform uniform frame sampling by 9.7 points on temporal grounding and by 17.3 points on JumpScore, while remaining competitive on long-video QA. The paper’s main claim is that perceptual token allocation should follow coding structure, not elapsed time, and it releases code, data, and model weights as open-source resources.
Original abstract
We introduce LLaVA-OneVision-2 (LLaVA-OV-2), the most capable vision-language model in the LLaVA-OneVision series to date, achieving superior performance across a broad range of multimodal benchmarks. The model builds on a native OneVision-Encoder and incorporates Windowed Attention for efficient local computation while maintaining native resolution. Its key advance is codec-stream tokenization: it treats compressed video as a continuous bit-cost stream, where bit-cost dynamics determine adaptive temporal groups, and motion-residual cues select salient spatial evidence into compact visual canvases. This allocation concentrates a limited token budget on event-bearing content, enabling more stable long-video token compression than fixed groups of pictures. A shared 3D RoPE further places codec canvases, sampled frames, and images in a unified spatiotemporal coordinate system. Furthermore, we build the LLaVA-OV-2 data and training stack around large-scale open supervision: approximately 8M re-captioned video samples for pretraining, a 4M-sample spatial corpus for fine-tuning. We also introduce JumpScore, a temporal-localization benchmark targeting fine-grained grounding in high-frequency, densely repeated motion, a regime underrepresented by existing video evaluations. A standout capability of LLaVA-OV-2 is its unified perception across video understanding, temporal grounding, spatial grounding, and manipulation-trace reasoning. On JumpScore, LLaVA-OneVision-2-8B reaches 74.9 JumpScore mAP, surpassing Qwen3-VL-8B (30.1) by +44.8 points; under matched visual-token budgets on the same benchmark, codec-stream inputs improve temporal grounding over frame sampling by +9.7 points. Across standard benchmarks, LLaVA-OneVision-2-8B further outperforms Qwen3-VL-8B by +4.3 average points on video tasks, +5.3 on spatial tasks, and +15.6 average J&F on tracking tasks.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.