Towards One-to-Many Temporal Grounding
AuthorsQi Xu, Yue Tan, Shihao Chen, Jiahao Meng, Anna Wang, Shunping Ji, Hao Fei, Jason Li
Resources
This paper tackles a harder version of video grounding where one text query can match multiple separate video segments, and it builds a benchmark plus training method to make multimodal models much better at finding all of them.
Key results
The first comprehensive OMTG benchmark contains 340 manually curated samples.
The instruction-tuning dataset is built from approximately 56k high-quality samples.
OMTG-4B achieves 43.65 EtF1 on OMTG Bench, the best reported score.
OMTG-4B outperforms Gemini 2.5 Pro by 15.85 points in EtF1 on OMTG Bench.
OMTG-4B outperforms Seed-1.8 by 15.61 points in EtF1 on OMTG Bench.
What the paper found
Towards One-to-Many Temporal Grounding, from a team including authors at Wuhan University, Peking University, Nanyang Technological University, and the National University of Singapore, formalizes a new video understanding problem where one text query must retrieve all disjoint matching moments in a video, rather than a single segment. The paper argues that existing MLLMs such as Gemini 2.5 Pro, Gemini 3 Pro, Seed-1.8, and Qwen3-VL fail because they lack event-cardinality perception, often merging multiple occurrences into one span or hallucinating extra segments. To measure this setting correctly, the authors introduce the OMTG Bench, a 340-sample benchmark with manually verified multi-segment annotations, plus new metrics: Count Accuracy, Temporal F1, and Effective Temporal F1, which zeroes out examples with wrong segment counts. They also build a 56k-sample instruction-tuning dataset using a five-stage pipeline based on Qwen3-VL-235B and Gemini 2.5 Pro, including repetitive-event discovery, grounding, strict visual veto filtering, recall checking, query refinement, and query-guided dense captioning. Training combines supervised fine-tuning with GRPO reinforcement learning using a composite reward that explicitly optimizes temporal IoU, count accuracy, caption quality, and output length. On OMTG Bench, their OMTG-4B model reaches 43.65 EtF1, outperforming Gemini 2.5 Pro by 15.85 points and Seed-1.8 by 15.61 points, while also improving standard one-to-one grounding on TimeLens and general video understanding on VideoMME.
Original abstract
Temporal Grounding (TG) aims to localize video segments corresponding to a textual query. Prior research predominantly focuses on single-segment retrieval. Real-world scenarios, however, often require localizing multiple disjoint segments for a single query -- a setting we term One-to-Many Temporal Grounding (OMTG). Previous state-of-the-art MLLMs, optimized for one-to-one settings, struggle in this context, often yielding near-zero scores due to a lack of event cardinality perception. To bridge this gap, we present a systematic solution with three key contributions. First, we establish the first comprehensive OMTG benchmark, introducing Count Accuracy (C-Acc) and Effective Temporal F1 (EtF1) as evaluation metrics. Second, we curate a high-quality OMTG dataset comprising 56k samples through a sophisticated construction pipeline. Third, we develop novel temporal and caption reward functions specifically designed for OMTG. In particular, the caption reward leverages Chain-of-Thought reasoning over dense video captions to explicitly guide policy optimization toward both preciseness and completeness. Extensive experiments show our model achieves a new state-of-the-art EtF1 of 43.65\% on OMTG Bench, outperforming Gemini 2.5 Pro and Seed-1.8 by 15.85\% and 15.61\%, respectively.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.