Accurate distance estimation for small drones in long-range imagery is important for tracking and situational awareness, yet remains challenging due to extreme target scale variation, background clutter, and noisy visual cues. This paper studies monocular drone distance estimation using image crops together with bounding-box geometry, a practical setting in which a detector provides a candidate drone region and the model predicts range from appearance and box-derived features. We evaluate a Droneranger-style baseline, and introduce a new DroneDAR (Drone Detection And Ranging) model that combines a convolutional backbone with explicit bounding-box cues through a lightweight gating mechanism. Experiments analyze how backbone capacity, crop resolution, and regression loss functions affect performance across distance regimes. We further examine common failure modes at long distances, including sensitivity to bounding-box noise and reduced texture detail in the crop. The results provide guidance for designing and training range estimators that remain robust under real-world long-range conditions and highlight directions for improving reliability when drones occupy only a few pixels.
翻译:远距离图像中小型无人机的精确距离估计对目标跟踪与态势感知至关重要,然而极端的目标尺度变化、背景杂波和视觉噪声使其依然充满挑战。本文研究了利用图像裁剪块及边界框几何特征进行单目无人机距离估计的方法(一种实用场景:检测器提供候选无人机区域,模型根据外观与框衍生特征预测距离)。我们评估了基于Droneranger的基线模型,并提出新型DroneDAR(无人机检测与测距)模型,该模型通过轻量级门控机制将卷积骨干网络与显式边界框线索相结合。实验分析了骨干网络容量、裁剪分辨率与回归损失函数在不同距离区间内对性能的影响。我们还进一步探究了远距离下的常见失效模式,包括对边界框噪声的敏感性及裁剪图像纹理细节的减少。研究结果为设计和训练在实际远距离条件下保持鲁棒的测距模型提供了指导,并指出了当无人机仅占据少量像素时提升可靠性的发展方向。