Vision-based Unmanned Aerial Vehicles (UAVs) frameworks aid human search tasks by detecting and recognizing specific individuals, then tracking and following them while maintaining a safe distance. A key safety requirement for UAV following is the accurate estimation of the distance between camera and target object under real-world conditions, achieved by fusing multiple image modalities. As part of the system for automatic people detection and face recognition using deep learning, in this paper we present the fusion of depth camera measurements and monocular camera-to-body distance estimation for robust tracking and following. Deep learning based filtering of depth camera data and estimation of camera-to-body distance from a monocular camera are achieved with YOLO-pose, enabling real-time fusion of depth information using the Extended Kalman Filter (EKF) algorithm. The proposed subsystem, designed for use in drones, estimates and measures the distance between the depth camera and the human body keypoints, to maintain the safe distance between the drone and the human target. Our system provides an accurate estimated distance, which has been validated against motion capture ground truth data. The system has been tested in real time indoors, where it reduces the average errors, RMSE and standard deviations of distance estimation up to 15,3% in three tested scenarios. Based on the test results, the EKF fusion-based approach increases the depth detection range by reducing the errors outside the optimal depth camera working range. It also shows improved robustness and precision in challenging conditions, such as reflections and poor visibility, making it suitable for SAR.
翻译:基于视觉的无人驾驶飞行器框架通过检测与识别特定个体,随后在保持安全距离的前提下对其进行跟踪与追随,从而辅助人类搜索任务。无人机追随的关键安全要求是在真实场景下,通过融合多种图像模态,精确估计相机与目标物体之间的距离。作为基于深度学习的自动人员检测与面部识别系统的一部分,本文提出将深度相机测量值与单目相机-人体距离估计相融合,以实现鲁棒的跟踪与追随。通过YOLO-pose实现基于深度学习的深度相机数据滤波和单目相机人体距离估计,并利用扩展卡尔曼滤波算法实现深度信息的实时融合。该子系统专为无人机设计,通过估计与测量深度相机至人体关键点的距离,维持无人机与人体目标之间的安全距离。我们的系统能够提供精确的距离估计,该估计值已通过运动捕捉真值数据验证。系统在室内实时测试中,在三个测试场景下将距离估计的平均误差、均方根误差和标准差降低了最多15.3%。测试结果表明,基于EKF融合的方法通过降低最优深度相机工作范围外的误差,提升了深度探测范围。该方法在反射、低能见度等挑战性条件下展现出更高的鲁棒性与精度,适用于搜救任务。