In recent years, several efforts have been aimed at improving the robustness of vision models to domains and environments unseen during training. An important practical problem pertains to models deployed in a new geography that is under-represented in the training dataset, posing a direct challenge to fair and inclusive computer vision. In this paper, we study the problem of geographic robustness and make three main contributions. First, we introduce a large-scale dataset GeoNet for geographic adaptation containing benchmarks across diverse tasks like scene recognition (GeoPlaces), image classification (GeoImNet) and universal adaptation (GeoUniDA). Second, we investigate the nature of distribution shifts typical to the problem of geographic adaptation and hypothesize that the major source of domain shifts arise from significant variations in scene context (context shift), object design (design shift) and label distribution (prior shift) across geographies. Third, we conduct an extensive evaluation of several state-of-the-art unsupervised domain adaptation algorithms and architectures on GeoNet, showing that they do not suffice for geographical adaptation, and that large-scale pre-training using large vision models also does not lead to geographic robustness. Our dataset is publicly available at https://tarun005.github.io/GeoNet.
翻译:近年来,多项研究致力于提升视觉模型对训练中未见领域和环境的鲁棒性。一个重要的实际问题涉及在训练数据集中代表性不足的新地域部署模型,这对公平包容的计算机视觉提出了直接挑战。本文研究地域鲁棒性问题,并做出三项主要贡献:首先,我们推出大规模数据集GeoNet,用于地理适应性研究,包含场景识别(GeoPlaces)、图像分类(GeoImNet)和通用适应性(GeoUniDA)等多样化任务的基准。其次,我们探究地理适应性问题的典型分布偏移特征,假设域偏移的主要来源是不同地域间场景上下文(上下文偏移)、物体设计(设计偏移)和标签分布(先验偏移)的显著差异。第三,我们在GeoNet上对多种最先进的无监督域适应算法和架构进行全面评估,结果表明它们不足以实现地理适应性,且大规模预训练视觉模型也无法带来地域鲁棒性。我们的数据集公开于 https://tarun005.github.io/GeoNet。