NTH

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs

AuthorsYuhan Zhu, Changlian Ma, Xiangyu Zeng, Xinhao Li, Zhiqiu Zhang, Songze Li, Jun Zhang, Tianxiang Jiang, Yuandong Yang, Ziang Yan, Zikang Wang, Xinyu Chen, Haoran Chen, Shaowei Zhang, Limin Wang

July 27, 2026 2 min read
Watch on YouTube
The one-line take

TimeLens2 helps video-language models identify all the moments supporting an answer, using better multi-span supervision and a reward designed for flexible temporal evidence sets.

Key results

93.2K
TimeLens2-93K grounding instances

Verified single- and multi-span supervision corpus

44.5
TimeLens2-2B average mIoU

Average across seven temporal-grounding benchmarks

47.7
TimeLens2-4B average mIoU

Average across seven temporal-grounding benchmarks

7.5
4B gain over Qwen3.5-397B-A17B

Average mIoU-point advantage

75.8%
Temporal Wasserstein zero-tIoU rescue

All-zero-tIoU GRPO groups receiving informative preferences

What the paper found

Researchers from Nanjing University and Shanghai AI Laboratory introduce TimeLens2, a generalist multimodal LLM for locating one or more evidence intervals in short, long, repetitive, interrogative, and egocentric videos. Its 93.2K-instance TimeLens2-93K corpus replaces brittle global annotation with a staged pipeline: Qwen3-VL-235B-A22B produces hierarchical timestamped captions, Kimi-K2.5 proposes queries and coarse spans, Qwen3-VL-30B-A3B and TimeLens-8B independently relocalize evidence, and consensus, Qwen3-VL-Embedding semantic verification, and local boundary refinement filter and sharpen labels. During GRPO optimization, a temporal Wasserstein reward computes exact one-dimensional W1 between uniform distributions over merged interval supports, complementing temporal IoU with dense, matching-free feedback for disjoint or differently fragmented predictions. Across seven benchmarks—covering CharadesTL, ActivityNetTL, QVHighlightsTL, VUE-TR, VUE-TR-V2, MomentSeeker, and Ego4D-NLQ—TimeLens2-2B reaches 44.5 average mIoU, while TimeLens2-4B and TimeLens2-8B reach 47.7 and 48.0; the variants improve over their Qwen3-VL backbones by 14.2, 13.0, and 18.1 points. The 4B model surpasses Qwen3.5-397B-A17B by 7.5 points on average, and the 8B model approaches Google’s Gemini 2.5 Pro. Wasserstein feedback rescues 75.8% of all-zero-tIoU GRPO groups, demonstrating that interval-set geometry, rather than timestamp syntax alone, drives the gains.

Original abstract

Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when the supporting evidence occurs. We study generalist video temporal grounding, in which one model predicts a variable-cardinality set of evidence intervals across video lengths, domains, query forms, and viewpoints. Existing training strategies are misaligned with this set-valued task: long-video labels often rely on brittle one-pass annotation, while reinforcement-learning rewards either fail to distinguish non-overlapping predictions or require fragile segment matching. TimeLens2 treats temporal evidence as an interval set throughout supervision and optimization. TimeLens2-93K constructs reliable multi-span supervision through caption-derived proposals, independent localization, cross-agent consensus, semantic verification, and boundary refinement. Our temporal Wasserstein reward computes exact one-dimensional \(W_1\) between uniform distributions over merged interval supports, providing dense, matching-free feedback under unequal cardinalities and equivalent fragmentation; temporal IoU complements it with precise-overlap feedback. Across seven benchmarks, TimeLens2-2B outperforms all size-matched baselines on every benchmark, while the 4B and 8B variants achieve state-of-the-art performance, surpassing open-source models with up to 397B parameters. The 2B, 4B, and 8B variants improve over their Qwen3-VL backbones by 14.2, 13.0, and 18.1 mIoU points, respectively.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis