In the realm of video analysis, the field of multiple object tracking (MOT) assumes paramount importance, with the motion state of objects-whether static or dynamic relative to the ground-holding practical significance across diverse scenarios. However, the extant literature exhibits a notable dearth in the exploration of this aspect. Deep learning methodologies encounter challenges in accurately discerning object motion states, while conventional approaches reliant on comprehensive mathematical modeling may yield suboptimal tracking accuracy. To address these challenges, we introduce a Model-Data-Driven Motion State Judgment Object Tracking Method (MoD2T). This innovative architecture adeptly amalgamates traditional mathematical modeling with deep learning-based multi-object tracking frameworks. The integration of mathematical modeling and deep learning within MoD2T enhances the precision of object motion state determination, thereby elevating tracking accuracy. Our empirical investigations comprehensively validate the efficacy of MoD2T across varied scenarios, encompassing unmanned aerial vehicle surveillance and street-level tracking. Furthermore, to gauge the method's adeptness in discerning object motion states, we introduce the Motion State Validation F1 (MVF1) metric. This novel performance metric aims to quantitatively assess the accuracy of motion state classification, furnishing a comprehensive evaluation of MoD2T's performance. Elaborate experimental validations corroborate the rationality of MVF1. In order to holistically appraise MoD2T's performance, we meticulously annotate several renowned datasets and subject MoD2T to stringent testing. Remarkably, under conditions characterized by minimal or moderate camera motion, the achieved MVF1 values are particularly noteworthy, with exemplars including 0.774 for the KITTI dataset, 0.521 for MOT17, and 0.827 for UAVDT.
翻译:在视频分析领域,多目标跟踪(MOT)具有至关重要的意义,而物体相对于地面的运动状态(静态或动态)在多种场景中具有实际应用价值。然而,现有文献对该方面的探索明显不足。深度学习方法在准确辨别物体运动状态方面面临挑战,而依赖全面数学建模的传统方法可能导致跟踪精度欠佳。为解决这些问题,我们提出了一种模型-数据驱动的运动状态判断目标跟踪方法(MoD2T)。该创新架构巧妙地将传统数学建模与基于深度学习的多目标跟踪框架相结合。MoD2T中数学建模与深度学习的融合提升了物体运动状态判断的精确性,从而提高了跟踪精度。我们的实证研究全面验证了MoD2T在多种场景(包括无人机监控和街道级跟踪)中的有效性。此外,为衡量该方法在辨别物体运动状态方面的能力,我们引入了运动状态验证F1值(MVF1)指标。这一新型性能指标旨在定量评估运动状态分类的准确性,为MoD2T的性能提供全面评价。详尽的实验验证证实了MVF1的合理性。为全面评估MoD2T的性能,我们细致标注了多个知名数据集,并对MoD2T进行了严格测试。值得注意的是,在相机运动较小或适中的条件下,所获得的MVF1值尤为突出,例如KITTI数据集为0.774,MOT17为0.521,UAVDT为0.827。