Traffic accident prediction in driving videos aims to provide an early warning of the accident occurrence, and supports the decision making of safe driving systems. Previous works usually concentrate on the spatial-temporal correlation of object-level context, while they do not fit the inherent long-tailed data distribution well and are vulnerable to severe environmental change. In this work, we propose a Cognitive Accident Prediction (CAP) method that explicitly leverages human-inspired cognition of text description on the visual observation and the driver attention to facilitate model training. In particular, the text description provides a dense semantic description guidance for the primary context of the traffic scene, while the driver attention provides a traction to focus on the critical region closely correlating with safe driving. CAP is formulated by an attentive text-to-vision shift fusion module, an attentive scene context transfer module, and the driver attention guided accident prediction module. We leverage the attention mechanism in these modules to explore the core semantic cues for accident prediction. In order to train CAP, we extend an existing self-collected DADA-2000 dataset (with annotated driver attention for each frame) with further factual text descriptions for the visual observations before the accidents. Besides, we construct a new large-scale benchmark consisting of 11,727 in-the-wild accident videos with over 2.19 million frames (named as CAP-DATA) together with labeled fact-effect-reason-introspection description and temporal accident frame label. Based on extensive experiments, the superiority of CAP is validated compared with state-of-the-art approaches. The code, CAP-DATA, and all results will be released in \url{https://github.com/JWFanggit/LOTVS-CAP}.
翻译:驾驶视频中的交通事故预测旨在为事故发生的早期预警提供支持,并辅助安全驾驶系统的决策。以往研究通常侧重于对象级上下文的时空关联,但其难以适应固有的长尾数据分布,且对剧烈环境变化的鲁棒性较弱。为此,本文提出一种认知事故预测(CAP)方法,该方法显式利用人类启发的视觉观测文本描述认知与驾驶员注意力来促进模型训练。具体而言,文本描述为交通场景的主要上下文提供密集的语义描述引导,而驾驶员注意力提供牵引力以聚焦与安全驾驶密切相关的关键区域。CAP由注意力驱动的文本-视觉移位融合模块、注意力驱动的场景上下文迁移模块及驾驶员注意力引导的事故预测模块构成。我们利用这些模块中的注意力机制来探索事故预测的核心语义线索。为训练CAP,我们在现有自收集的DADA-2000数据集(每帧标注驾驶员注意力)基础上,进一步为事故前的视觉观测添加事实性文本描述。此外,我们构建了一个包含11,727个野外事故视频(超过219万帧)的大规模新基准(命名为CAP-DATA),并标注了事实-效果-原因-内省描述及时间维度的事故帧标签。大量实验验证了CAP相较于现有最优方法的优越性。代码、CAP-DATA及所有结果将在\url{https://github.com/JWFanggit/LOTVS-CAP}开放。