NTH

Towards One-to-Many Temporal Grounding

AuthorsQi Xu, Yue Tan, Shihao Chen, Jiahao Meng, Anna Wang, Shunping Ji, Hao Fei, Jason Li

June 7, 2026 2 min read
Watch on YouTube
The one-line take

This paper tackles a harder version of video grounding where one text query can match multiple separate video segments, and it builds a benchmark plus training method to make multimodal models much better at finding all of them.

Key results

340
OMTG Bench

The first comprehensive OMTG benchmark contains 340 manually curated samples.

56k
OMTG dataset

The instruction-tuning dataset is built from approximately 56k high-quality samples.

43.65
OMTG-4B EtF1

OMTG-4B achieves 43.65 EtF1 on OMTG Bench, the best reported score.

15.85
Gemini 2.5 Pro EtF1 gap

OMTG-4B outperforms Gemini 2.5 Pro by 15.85 points in EtF1 on OMTG Bench.

15.61
Seed-1.8 EtF1 gap

OMTG-4B outperforms Seed-1.8 by 15.61 points in EtF1 on OMTG Bench.

What the paper found

Towards One-to-Many Temporal Grounding, from a team including authors at Wuhan University, Peking University, Nanyang Technological University, and the National University of Singapore, formalizes a new video understanding problem where one text query must retrieve all disjoint matching moments in a video, rather than a single segment. The paper argues that existing MLLMs such as Gemini 2.5 Pro, Gemini 3 Pro, Seed-1.8, and Qwen3-VL fail because they lack event-cardinality perception, often merging multiple occurrences into one span or hallucinating extra segments. To measure this setting correctly, the authors introduce the OMTG Bench, a 340-sample benchmark with manually verified multi-segment annotations, plus new metrics: Count Accuracy, Temporal F1, and Effective Temporal F1, which zeroes out examples with wrong segment counts. They also build a 56k-sample instruction-tuning dataset using a five-stage pipeline based on Qwen3-VL-235B and Gemini 2.5 Pro, including repetitive-event discovery, grounding, strict visual veto filtering, recall checking, query refinement, and query-guided dense captioning. Training combines supervised fine-tuning with GRPO reinforcement learning using a composite reward that explicitly optimizes temporal IoU, count accuracy, caption quality, and output length. On OMTG Bench, their OMTG-4B model reaches 43.65 EtF1, outperforming Gemini 2.5 Pro by 15.85 points and Seed-1.8 by 15.61 points, while also improving standard one-to-one grounding on TimeLens and general video understanding on VideoMME.

Original abstract

Temporal Grounding (TG) aims to localize video segments corresponding to a textual query. Prior research predominantly focuses on single-segment retrieval. Real-world scenarios, however, often require localizing multiple disjoint segments for a single query -- a setting we term One-to-Many Temporal Grounding (OMTG). Previous state-of-the-art MLLMs, optimized for one-to-one settings, struggle in this context, often yielding near-zero scores due to a lack of event cardinality perception. To bridge this gap, we present a systematic solution with three key contributions. First, we establish the first comprehensive OMTG benchmark, introducing Count Accuracy (C-Acc) and Effective Temporal F1 (EtF1) as evaluation metrics. Second, we curate a high-quality OMTG dataset comprising 56k samples through a sophisticated construction pipeline. Third, we develop novel temporal and caption reward functions specifically designed for OMTG. In particular, the caption reward leverages Chain-of-Thought reasoning over dense video captions to explicitly guide policy optimization toward both preciseness and completeness. Extensive experiments show our model achieves a new state-of-the-art EtF1 of 43.65\% on OMTG Bench, outperforming Gemini 2.5 Pro and Seed-1.8 by 15.85\% and 15.61\%, respectively.

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis