Although there is a significant development in 3D Multi-view Multi-person Tracking (3D MM-Tracking), current 3D MM-Tracking frameworks are designed separately for footprint and pose tracking. Specifically, frameworks designed for footprint tracking cannot be utilized in 3D pose tracking, because they directly obtain 3D positions on the ground plane with a homography projection, which is inapplicable to 3D poses above the ground. In contrast, frameworks designed for pose tracking generally isolate multi-view and multi-frame associations and may not be robust to footprint tracking, since footprint tracking utilizes fewer key points than pose tracking, which weakens multi-view association cues in a single frame. This study presents a Unified Multi-view Multi-person Tracking framework to bridge the gap between footprint tracking and pose tracking. Without additional modifications, the framework can adopt monocular 2D bounding boxes and 2D poses as the input to produce robust 3D trajectories for multiple persons. Importantly, multi-frame and multi-view information are jointly employed to improve the performance of association and triangulation. The effectiveness of our framework is verified by accomplishing state-of-the-art performance on the Campus and Shelf datasets for 3D pose tracking, and by comparable results on the WILDTRACK and MMPTRACK datasets for 3D footprint tracking.
翻译:尽管三维多视角多人追踪(3D MM-Tracking)领域取得了显著进展,但当前的三维多视角多人追踪框架分别针对足迹追踪和姿态追踪独立设计。具体而言,针对足迹追踪设计的框架无法用于三维姿态追踪,因为它们通过单应投影直接获取地面平面上的三维位置,而这不适用于地面以上三维姿态的定位。相反,针对姿态追踪设计的框架通常将多视角关联与多帧关联相分离,可能对足迹追踪缺乏鲁棒性,因为足迹追踪使用的关键点数量少于姿态追踪,从而削弱了单帧中的多视角关联线索。本研究提出了一种统一的多视角多人追踪框架,旨在弥合足迹追踪与姿态追踪之间的差距。无需额外修改,该框架即可采用单目二维边界框和二维姿态作为输入,为多人生成鲁棒的三维轨迹。关键之处在于,多帧与多视角信息被联合利用,以提升关联与三角测量的性能。通过在Campus和Shelf数据集上实现三维姿态追踪的最优性能,并在WILDTRACK和MMPTRACK数据集上获得三维足迹追踪的同等结果,验证了我们框架的有效性。