Conformal Coverage Guarantees for Any Video Temporal Grounder
AuthorsAseel Mohamed, Rasul Khanbayov, Erchin Serpedin, Hasan Kurban
Resources
COVER turns video moment predictions into calibrated time regions that say not only where an event is, but how reliably it is covered.
Key results
Realized coverage at a requested target of 0.90 for ActivityNet-Captions.
SRAM’s evidential uncertainty region achieved 0.660 against its nominal 0.80 target.
Coverage varied by 0.904 across hand-picked fixed margins, from 0.096 to 1.000.
What the paper found
Cover is a post-hoc, model-agnostic wrapper for video temporal grounding: instead of returning one unqualified interval, it calibrates a temporal nonconformity score on held-out labeled examples and widens each prediction into a region with finite-sample, distribution-free coverage of at least 1 − α under exchangeability. It requires no retraining or white-box access, and works with trained localizers such as QD-DETR and black-box video-language models such as Qwen2.5-VL-7B-Instruct. The method supports a length-normalized two-sided boundary score, a separate per-boundary calibration for asymmetric errors, and a super-level-set score that can produce disconnected relevance regions. Across Charades-STA, ActivityNet-Captions, and QVHighlights, covering three benchmarks and five grounders, realized coverage closely follows the requested target: for example, ActivityNet-Captions reaches 0.895 at a 0.90 target. Cover also exposes why learned uncertainty is not automatically reliable: the evidential SRAM grounder achieves only 0.660 against its nominal 0.80 target, whereas Cover restores calibrated coverage on the same predictions. Fixed, hand-picked margins are unstable, producing coverage from 0.096 to 1.000, a 0.904 range across models and datasets. Calibration also reveals that Qwen’s boundary errors are concentrated at event offsets, while its onsets are nearly accurate. The main limitation is that guarantees are marginal and depend on exchangeability; dataset shift, within-video dependence, or conditioning on difficult subgroups can reduce validity, motivating weighted or Mondrian calibration.
Original abstract
Event boundaries in continuous video are ambiguous: re-annotate the same query-video pair and independent annotators mark moments that overlap by less than half on a large fraction of samples. The ground truth for video temporal grounding is therefore a distribution over intervals, yet every grounder returns a single interval with no statement of reliability, so at deployment a wrong interval is indistinguishable from a right one. COVER changes the output object: a post-hoc, model-agnostic wrapper that turns any grounder, a trained localizer or a black-box video--language model, into one that emits a temporal region containing the true moment with probability at least $1-α$, by calibrating the quantile of a temporal nonconformity score on held-out labels and widening the base prediction by that amount. The guarantee is finite-sample and distribution-free under exchangeability, and requires neither retraining nor white-box access. We give two score families, a two-sided boundary-widening score for grounders that emit an interval and a super-level-set score for grounders that emit a relevance signal, and develop theory specific to grounding that bounds how large the certified region becomes, when coverage survives conditioning on event length, and how it degrades when moments from one video break exchangeability. Across three benchmarks and five grounders, realized coverage tracks the target, and calibration exposes what point metrics hide.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.