In this study, we introduce DeepLocalization, an innovative framework devised for the real-time localization of actions tailored explicitly for monitoring driver behavior. Utilizing the power of advanced deep learning methodologies, our objective is to tackle the critical issue of distracted driving-a significant factor contributing to road accidents. Our strategy employs a dual approach: leveraging Graph-Based Change-Point Detection for pinpointing actions in time alongside a Video Large Language Model (Video-LLM) for precisely categorizing activities. Through careful prompt engineering, we customize the Video-LLM to adeptly handle driving activities' nuances, ensuring its classification efficacy even with sparse data. Engineered to be lightweight, our framework is optimized for consumer-grade GPUs, making it vastly applicable in practical scenarios. We subjected our method to rigorous testing on the SynDD2 dataset, a complex benchmark for distracted driving behaviors, where it demonstrated commendable performance-achieving 57.5% accuracy in event classification and 51% in event detection. These outcomes underscore the substantial promise of DeepLocalization in accurately identifying diverse driver behaviors and their temporal occurrences, all within the bounds of limited computational resources.
翻译:本研究提出DeepLocalization,这是一种专为实时监测驾驶员行为而设计的动作定位创新框架。借助先进深度学习方法的优势,我们旨在解决分心驾驶这一导致交通事故的关键因素。本策略采用双轨方法:利用基于图的变点检测技术实现动作时序定位,同时结合视频大语言模型(Video-LLM)进行精准行为分类。通过精心设计的提示工程,我们定制化视频大语言模型以精准适应驾驶行为的细微特征,确保其在数据稀疏情况下仍具备高效分类能力。该框架经轻量化设计后,可在消费级GPU上高效运行,显著提升了实际场景的适用性。我们在分心驾驶行为复杂基准数据集SynDD2上进行了严格测试,该模型在事件分类任务中取得57.5%的准确率,事件检测准确率达51%。这些成果充分表明,DeepLocalization在有限计算资源条件下,能够精准识别多样化驾驶行为及其时间分布,展现出巨大的应用潜力。