Aerial-ground person re-identification (Re-ID) presents unique challenges in computer vision, stemming from the distinct differences in viewpoints, poses, and resolutions between high-altitude aerial and ground-based cameras. Existing research predominantly focuses on ground-to-ground matching, with aerial matching less explored due to a dearth of comprehensive datasets. To address this, we introduce AG-ReID.v2, a dataset specifically designed for person Re-ID in mixed aerial and ground scenarios. This dataset comprises 100,502 images of 1,615 unique individuals, each annotated with matching IDs and 15 soft attribute labels. Data were collected from diverse perspectives using a UAV, stationary CCTV, and smart glasses-integrated camera, providing a rich variety of intra-identity variations. Additionally, we have developed an explainable attention network tailored for this dataset. This network features a three-stream architecture that efficiently processes pairwise image distances, emphasizes key top-down features, and adapts to variations in appearance due to altitude differences. Comparative evaluations demonstrate the superiority of our approach over existing baselines. We plan to release the dataset and algorithm source code publicly, aiming to advance research in this specialized field of computer vision. For access, please visit https://github.com/huynguyen792/AG-ReID.v2.
翻译:空中-地面行人重识别(Re-ID)在计算机视觉领域面临独特挑战,源于高空航拍相机与地面相机之间在视角、姿态和分辨率上的显著差异。现有研究主要集中于地-地匹配,而由于缺乏全面数据集,航拍匹配研究相对较少。为解决此问题,我们提出了AG-ReID.v2——一个专为混合空中与地面场景行人重识别设计的数据集。该数据集包含来自1,615个不同个体的100,502张图像,每张图像均标注匹配ID和15个软属性标签。数据通过无人机、固定闭路电视及集成智能眼镜的相机从多视角采集,提供了丰富的个体内变化。此外,我们针对该数据集开发了一种可解释注意力网络。该网络采用三流架构,高效处理成对图像距离、强调关键自上而下特征,并适应海拔差异导致的外观变化。对比评估表明,我们的方法优于现有基线。我们计划公开发布该数据集与算法源代码,旨在推动计算机视觉这一专业领域的研究进展。如需访问,请访问https://github.com/huynguyen792/AG-ReID.v2。