NTH

Conformal Coverage Guarantees for Any Video Temporal Grounder

AuthorsAseel Mohamed, Rasul Khanbayov, Erchin Serpedin, Hasan Kurban

August 17, 2026 2 min read
Watch on YouTube
The one-line take

COVER turns video moment predictions into calibrated time regions that say not only where an event is, but how reliably it is covered.

Key results

0.895
ActivityNet coverage

Realized coverage at a requested target of 0.90 for ActivityNet-Captions.

0.660
SRAM nominal coverage failure

SRAM’s evidential uncertainty region achieved 0.660 against its nominal 0.80 target.

0.904
Fixed-margin coverage range

Coverage varied by 0.904 across hand-picked fixed margins, from 0.096 to 1.000.

What the paper found

Cover is a post-hoc, model-agnostic wrapper for video temporal grounding: instead of returning one unqualified interval, it calibrates a temporal nonconformity score on held-out labeled examples and widens each prediction into a region with finite-sample, distribution-free coverage of at least 1 − α under exchangeability. It requires no retraining or white-box access, and works with trained localizers such as QD-DETR and black-box video-language models such as Qwen2.5-VL-7B-Instruct. The method supports a length-normalized two-sided boundary score, a separate per-boundary calibration for asymmetric errors, and a super-level-set score that can produce disconnected relevance regions. Across Charades-STA, ActivityNet-Captions, and QVHighlights, covering three benchmarks and five grounders, realized coverage closely follows the requested target: for example, ActivityNet-Captions reaches 0.895 at a 0.90 target. Cover also exposes why learned uncertainty is not automatically reliable: the evidential SRAM grounder achieves only 0.660 against its nominal 0.80 target, whereas Cover restores calibrated coverage on the same predictions. Fixed, hand-picked margins are unstable, producing coverage from 0.096 to 1.000, a 0.904 range across models and datasets. Calibration also reveals that Qwen’s boundary errors are concentrated at event offsets, while its onsets are nearly accurate. The main limitation is that guarantees are marginal and depend on exchangeability; dataset shift, within-video dependence, or conditioning on difficult subgroups can reduce validity, motivating weighted or Mondrian calibration.

Original abstract

Event boundaries in continuous video are ambiguous: re-annotate the same query-video pair and independent annotators mark moments that overlap by less than half on a large fraction of samples. The ground truth for video temporal grounding is therefore a distribution over intervals, yet every grounder returns a single interval with no statement of reliability, so at deployment a wrong interval is indistinguishable from a right one. COVER changes the output object: a post-hoc, model-agnostic wrapper that turns any grounder, a trained localizer or a black-box video--language model, into one that emits a temporal region containing the true moment with probability at least $1-α$, by calibrating the quantile of a temporal nonconformity score on held-out labels and widening the base prediction by that amount. The guarantee is finite-sample and distribution-free under exchangeability, and requires neither retraining nor white-box access. We give two score families, a two-sided boundary-widening score for grounders that emit an interval and a super-level-set score for grounders that emit a relevance signal, and develop theory specific to grounding that bounds how large the certified region becomes, when coverage survives conditioning on event length, and how it degrades when moments from one video break exchangeability. Across three benchmarks and five grounders, realized coverage tracks the target, and calibration exposes what point metrics hide.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis