Unsupervised Domain Adaptive Object Detection (UDA-OD) uses unlabelled data to improve the reliability of robotic vision systems in open-world environments. Previous approaches to UDA-OD based on self-training have been effective in overcoming changes in the general appearance of images. However, shifts in a robot's deployment environment can also impact the likelihood that different objects will occur, termed class distribution shift. Motivated by this, we propose a framework for explicitly addressing class distribution shift to improve pseudo-label reliability in self-training. Our approach uses the domain invariance and contextual understanding of a pre-trained joint vision and language model to predict the class distribution of unlabelled data. By aligning the class distribution of pseudo-labels with this prediction, we provide weak supervision of pseudo-label accuracy. To further account for low quality pseudo-labels early in self-training, we propose an approach to dynamically adjust the number of pseudo-labels per image based on model confidence. Our method outperforms state-of-the-art approaches on several benchmarks, including a 4.7 mAP improvement when facing challenging class distribution shift.
翻译:无监督域自适应目标检测(UDA-OD)利用无标注数据来提升机器视觉系统在开放环境中的可靠性。基于自训练的现有UDA-OD方法在应对图像整体外观变化方面表现有效。然而,机器人部署环境的偏移也可能影响不同目标出现的概率,即所谓的类分布偏移。受此启发,我们提出了一种显式处理类分布偏移的框架,以提升自训练中伪标签的可靠性。该方法利用预训练的视觉-语言联合模型的域不变性和上下文理解能力,预测无标注数据的类分布。通过将伪标签的类分布与该预测对齐,为伪标签的准确性提供弱监督。为进一步解决自训练初期伪标签质量低的问题,我们提出了一种基于模型置信度动态调整每张图像伪标签数量的方法。我们的方法在多个基准测试中优于现有最先进方法,在面对具有挑战性的类分布偏移时,平均精度均值(mAP)提升了4.7个百分点。