Multimodal large language models have made rapid progress in video temporal grounding, yet real-world applications routinely require localizing every event that satisfies compositional temporal and spatial conditions. Existing benchmarks fall short: they localize only a single moment per query, count without temporal conditions, or treat grounding and counting as disjoint tasks. We introduce CoMET-Bench for Conditional Multi-Event Temporal Grounding in long-form video, comprising 2789 queries over 600 videos averaging 33.8 minutes across five real-world domains, with each query composed from 4 temporal conditions, 3 spatial conditions, and a dedicated negative-query subset. We further propose a unified evaluation protocol jointly measuring counting, grounding, and negative-query recognition, including a new Rejection-F1 metric that prevents trivial gaming by lazy "always-empty" models. Benchmarking a broad suite of MLLMs, agent-based, and grounding-specialized methods reveals that existing approaches remain far from solving this task. Building on these findings, we propose CoMET-Agent, a training-free agentic framework that reformulates the task as structured search-and-aggregate, improving [email protected] by 6.1% over GPT-5 purely through structural reasoning. Failure analysis further surfaces three open directions: fine-grained entity tracking, position-uniform retrieval, and causal event pairing.
翻译:多模态大语言模型在视频时间定位领域取得了快速进展,然而实际应用通常需要定位每个满足组合时间和空间条件的事件。现有基准存在不足:它们仅能定位每个查询中的单个时刻,计数时不考虑时间条件,或者将定位和计数视为独立任务。我们提出了面向长视频的条件多事件时间定位基准CoMET-Bench,包含来自五个真实领域、平均时长33.8分钟的600个视频的2789个查询,每个查询由4个时间条件、3个空间条件和一个专门的负查询子集组成。我们进一步提出了一种统一评估协议,共同衡量计数、定位和负查询识别能力,包括一种新的Rejection-F1指标,以防止懒惰的“始终为空”模型通过取巧方式获胜。对多种MLLM、基于智能体以及专门定位方法的基准测试表明,现有方法远未解决此任务。基于这些发现,我们提出了CoMET-Agent,一个无需训练的智能体框架,将任务重新定义为结构化搜索与聚合,通过纯粹的推理结构将[email protected]指标提升了6.1%(优于GPT-5)。失败分析进一步揭示了三个开放方向:细粒度实体追踪、位置均匀检索和因果事件配对。