Multimodal large language models (MLLMs) have demonstrated impressive general competence in video understanding, yet their reliability for real-world Video Anomaly Detection (VAD) remains largely unexplored. Unlike conventional pipelines relying on reconstruction or pose-based cues, MLLMs enable a paradigm shift: treating anomaly detection as a language-guided reasoning task. In this work, we systematically evaluate state-of-the-art MLLMs on the ShanghaiTech and CHAD benchmarks by reformulating VAD as a binary classification task under weak temporal supervision. We investigate how prompt specificity and temporal window lengths (1s--3s) influence performance, focusing on the precision--recall trade-off. Our findings reveal a pronounced conservative bias in zero-shot settings; while models exhibit high confidence, they disproportionately favor the 'normal' class, resulting in high precision but a recall collapse that limits practical utility. We demonstrate that class-specific instructions can significantly shift this decision boundary, improving the peak F1-score on ShanghaiTech from 0.09 to 0.64, yet recall remains a critical bottleneck. These results highlight a significant performance gap for MLLMs in noisy environments and provide a foundation for future work in recall-oriented prompting and model calibration for open-world surveillance, which demands complex video understanding and reasoning.
翻译:多模态大语言模型(MLLMs)在视频理解方面展现出令人瞩目的通用能力,但其在实际视频异常检测(VAD)中的可靠性仍鲜有探索。与依赖于重建或姿态线索的传统流程不同,MLLMs开创了一种范式转换:将异常检测视为语言引导的推理任务。在本研究中,我们通过将VAD重新表述为弱时间监督下的二分类任务,系统评估了最先进的MLLMs在ShanghaiTech和CHAD基准上的表现。我们探究了提示特异性及时间窗口长度(1秒至3秒)对性能的影响,重点关注精确率-召回率的权衡。研究结果揭示了零样本设置中显著的保守偏差:尽管模型表现出高置信度,但它们过度偏向"正常"类别,导致高精确率但召回率骤降,从而限制了实际应用价值。我们证明,类别特异性指令能够显著改变这一决策边界,将ShanghaiBench上的峰值F1分数从0.09提升至0.64,但召回率仍是关键瓶颈。这些发现凸显了MLLMs在噪声环境下的显著性能差距,并为面向开放世界监控的召回导向提示工程与模型校准工作奠定了基础——这类监控场景需要复杂的视频理解与推理能力。