Multi-view multi-human association and tracking (MvMHAT), is a new but important problem for multi-person scene video surveillance, aiming to track a group of people over time in each view, as well as to identify the same person across different views at the same time, which is different from previous MOT and multi-camera MOT tasks only considering the over-time human tracking. This way, the videos for MvMHAT require more complex annotations while containing more information for self learning. In this work, we tackle this problem with a self-supervised learning aware end-to-end network. Specifically, we propose to take advantage of the spatial-temporal self-consistency rationale by considering three properties of reflexivity, symmetry and transitivity. Besides the reflexivity property that naturally holds, we design the self-supervised learning losses based on the properties of symmetry and transitivity, for both appearance feature learning and assignment matrix optimization, to associate the multiple humans over time and across views. Furthermore, to promote the research on MvMHAT, we build two new large-scale benchmarks for the network training and testing of different algorithms. Extensive experiments on the proposed benchmarks verify the effectiveness of our method. We have released the benchmark and code to the public.
翻译:多视角多人关联与追踪(MvMHAT)是多场景视频监控领域一个新颖且重要的问题,旨在同时实现每个视角中随时间推移的人群追踪,以及同一时刻跨不同视角的同一人员识别。这与传统仅关注跨时间人员追踪的多目标追踪(MOT)和多摄像机MOT任务有所不同。因此,MvMHAT视频在包含更丰富自学习信息的同时,需要更复杂的标注。在本工作中,我们通过一种自监督学习感知的端到端网络解决该问题。具体而言,我们提出利用时空自一致性原理,综合考虑自反性、对称性和传递性三种性质。除天然成立的自反性外,我们基于对称性与传递性设计了自监督学习损失函数,分别应用于外观特征学习和分配矩阵优化,以实现多人跨时间与跨视角的关联。此外,为促进MvMHAT研究,我们构建了两个大规模基准数据集,用于不同算法的网络训练与测试。在提出基准上的大量实验验证了本方法的有效性。我们已公开发布基准数据集和代码。