Surveillance videos are an essential component of daily life with various critical applications, particularly in public security. However, current surveillance video tasks mainly focus on classifying and localizing anomalous events. Existing methods are limited to detecting and classifying the predefined events with unsatisfactory semantic understanding, although they have obtained considerable performance. To address this issue, we propose a new research direction of surveillance video-and-language understanding, and construct the first multimodal surveillance video dataset. We manually annotate the real-world surveillance dataset UCF-Crime with fine-grained event content and timing. Our newly annotated dataset, UCA (UCF-Crime Annotation), contains 23,542 sentences, with an average length of 20 words, and its annotated videos are as long as 110.7 hours. Furthermore, we benchmark SOTA models for four multimodal tasks on this newly created dataset, which serve as new baselines for surveillance video-and-language understanding. Through our experiments, we find that mainstream models used in previously publicly available datasets perform poorly on surveillance video, which demonstrates the new challenges in surveillance video-and-language understanding. To validate the effectiveness of our UCA, we conducted experiments on multimodal anomaly detection. The results demonstrate that our multimodal surveillance learning can improve the performance of conventional anomaly detection tasks. All the experiments highlight the necessity of constructing this dataset to advance surveillance AI. The link to our dataset is provided at: https://xuange923.github.io/Surveillance-Video-Understanding.
翻译:监控视频是日常生活中的重要组成部分,具有多种关键应用,尤其在公共安全领域。然而,当前监控视频任务主要聚焦于异常事件的分类与定位。现有方法虽已取得可观性能,但仍局限于检测和分类预定义事件,语义理解能力不足。为解决这一问题,我们提出了监控视频与语言理解这一新的研究方向,并构建了首个多模态监控视频数据集。我们针对真实场景监控数据集UCF-Crime,手动标注了细粒度的事件内容与时间信息。新标注的数据集UCA(UCF-Crime Annotation)包含23,542条语句,平均长度为20个词,标注视频总时长达到110.7小时。此外,我们在此新数据集上对四种多模态任务的SOTA模型进行了基准测试,为监控视频与语言理解提供了新的基线方法。通过实验发现,此前公开数据集上表现优异的主流模型在监控视频任务中性能较差,这揭示了监控视频与语言理解面临的新挑战。为验证UCA的有效性,我们开展了多模态异常检测实验。结果表明,我们的多模态监控学习能够提升传统异常检测任务的性能。所有实验均凸显了构建该数据集对推进监控AI研究的必要性。数据集链接见:https://xuange923.github.io/Surveillance-Video-Understanding。