Video-based gait recognition has achieved impressive results in constrained scenarios. However, visual cameras neglect human 3D structure information, which limits the feasibility of gait recognition in the 3D wild world. Instead of extracting gait features from images, this work explores precise 3D gait features from point clouds and proposes a simple yet efficient 3D gait recognition framework, termed LidarGait. Our proposed approach projects sparse point clouds into depth maps to learn the representations with 3D geometry information, which outperforms existing point-wise and camera-based methods by a significant margin. Due to the lack of point cloud datasets, we built the first large-scale LiDAR-based gait recognition dataset, SUSTech1K, collected by a LiDAR sensor and an RGB camera. The dataset contains 25,239 sequences from 1,050 subjects and covers many variations, including visibility, views, occlusions, clothing, carrying, and scenes. Extensive experiments show that (1) 3D structure information serves as a significant feature for gait recognition. (2) LidarGait outperforms existing point-based and silhouette-based methods by a significant margin, while it also offers stable cross-view results. (3) The LiDAR sensor is superior to the RGB camera for gait recognition in the outdoor environment. The source code and dataset have been made available at https://lidargait.github.io.
翻译:基于视频的步态识别在受控场景中已取得显著成果。然而,视觉相机忽略了人体的三维结构信息,这限制了步态识别在三维现实世界中的可行性。本文探索从点云中提取精确的三维步态特征,而非从图像中提取步态特征,并提出一个简单高效的三维步态识别框架,称为LidarGait。该方法将稀疏点云投影为深度图,以学习具有三维几何信息的表征,其性能显著优于现有逐点方法和基于相机的方法。由于缺乏点云数据集,我们构建了首个大规模基于LiDAR的步态识别数据集SUSTech1K,通过LiDAR传感器和RGB相机采集。该数据集包含来自1050名受试者的25239个序列,涵盖多种变化因素,包括可见性、视角、遮挡、着装、携带物和场景。大量实验表明:(1)三维结构信息是步态识别的重要特征;(2)LidarGait在性能上显著优于现有基于点和基于轮廓的方法,同时提供稳定的跨视角结果;(3)在室外环境中,LiDAR传感器在步态识别方面优于RGB相机。源代码和数据集已发布于https://lidargait.github.io。