Robotic Perception in diverse domains such as low-light scenarios, where new modalities like thermal imaging and specialized night-vision sensors are increasingly employed, remains a challenge. Largely, this is due to the limited availability of labeled data. Existing Domain Adaptation (DA) techniques, while promising to leverage labels from existing well-lit RGB images, fail to consider the characteristics of the source domain itself. We holistically account for this factor by proposing Source Preparation (SP), a method to mitigate source domain biases. Our Almost Unsupervised Domain Adaptation (AUDA) framework, a label-efficient semi-supervised approach for robotic scenarios -- employs Source Preparation (SP), Unsupervised Domain Adaptation (UDA) and Supervised Alignment (SA) from limited labeled data. We introduce CityIntensified, a novel dataset comprising temporally aligned image pairs captured from a high-sensitivity camera and an intensifier camera for semantic segmentation and object detection in low-light settings. We demonstrate the effectiveness of our method in semantic segmentation, with experiments showing that SP enhances UDA across a range of visual domains, with improvements up to 40.64% in mIoU over baseline, while making target models more robust to real-world shifts within the target domain. We show that AUDA is a label-efficient framework for effective DA, significantly improving target domain performance with only tens of labeled samples from the target domain.
翻译:在诸如低光照场景等多样化领域中,机器人感知仍面临挑战,这些场景中越来越多地采用热成像和专用夜视传感器等新模态。这主要归因于标记数据的有限可用性。现有的域适应技术虽有潜力利用现有良好照明RGB图像的标签,但未能考虑源域本身的特性。我们通过提出源准备方法——一种缓解源域偏差的方法——来整体考虑这一因素。我们的几乎无监督域适应框架——一种适用于机器人场景的标签高效半监督方法——结合了源准备、无监督域适应和基于有限标记数据的监督对齐。我们引入了CityIntensified新型数据集,该数据集包含从高灵敏度相机和增强相机获取的时间对齐图像对,用于低光照环境下的语义分割和目标检测。我们通过实验证明了该方法在语义分割中的有效性,结果显示源准备可增强跨多种视觉域的无监督域适应性能,平均交并比相比基线提升高达40.64%,同时使目标模型对目标域内的实际偏移更具鲁棒性。我们证明几乎无监督域适应是一种用于有效域适应的标签高效框架,仅需目标域数十个标记样本即可显著提升目标域性能。