Video multimodal large language models (MLLMs) have made rapid progress on general and long-form video understanding, yet their ability to preserve brief answer-critical visual evidence remains underexplored. Many practical questions are determined by momentary visual events: localized actions or state transitions that may last only a few frames. Such evidence can be skipped by sparse frame sampling, suppressed by visual-token compression, or diluted by coarse temporal aggregation, causing failures that language-side reasoning cannot reliably recover. We introduce Moment-Video, a benchmark for diagnosing the temporal fidelity of video MLLMs through momentary visual event understanding. Each question is grounded in a localized, visually observable, and sampling-sensitive event, requiring models to notice, count, describe, or reason about transient evidence rather than rely on persistent objects, global scene context, or language priors. Moment-Video contains 1,000 human-verified video-QA pairs across 7 domains and 25 fine-grained subcategories, covering four task types: Temporal Occurrence, Temporal Counting, Action Description, and Temporal Reasoning. We evaluate 33 proprietary and open-source MLLMs on Moment-Video. The best-performing model, Seed-2.0-Pro, achieves only 39.6% overall accuracy, while most open-source models remain below 25%, revealing a substantial gap in momentary visual event understanding. Diagnostic analyses show that denser frame sampling improves some models but does not eliminate the bottleneck, and longer videos introduce stronger temporal-localization challenges. These findings suggest that current video MLLMs still lack temporally faithful representations for capturing, preserving, and using brief but decisive visual evidence.
翻译:视频多模态大语言模型(MLLMs)在通用和长视频理解方面取得了快速进展,但其保留简短的关键视觉证据的能力仍未被充分探索。许多实际问题由瞬时视觉事件决定:可能仅持续几帧的局部化动作或状态转换。这类证据可能因稀疏帧采样而遗漏、因视觉标记压缩而抑制,或因粗粒度时间聚合而稀释,导致语言端推理无法可靠恢复这些失效。我们提出时刻视频,一个通过瞬时视觉事件理解来诊断视频MLLMs时间保真度的基准。每个问题都基于局部化、视觉可观察且对采样敏感的事件,要求模型关注、计数、描述或推理瞬时证据,而非依赖持久物体、全局场景背景或语言先验。时刻视频包含1000个人工验证的视频问答对,覆盖7个领域和25个细分子类别,涵盖四种任务类型:时间发生、时间计数、动作描述和时间推理。我们在时刻视频上评估了33个专有和开源MLLMs。表现最佳模型Seed-2.0-Pro仅达到39.6%的总体准确率,而多数开源模型低于25%,揭示了在瞬时视觉事件理解上的显著差距。诊断分析表明,更密集的帧采样改善了一些模型但未消除瓶颈,而更长视频引入了更强的时间定位挑战。这些发现表明,当前视频MLLMs仍缺乏捕捉、保留和利用简短但决定性视觉证据的时间保真表示。