Generalized Category Discovery (GCD) is a challenging task in which, given a partially labelled dataset, models must categorize all unlabelled instances, regardless of whether they come from labelled categories or from new ones. In this paper, we challenge a remaining assumption in this task: that all images share the same domain. Specifically, we introduce a new task and method to handle GCD when the unlabelled data also contains images from different domains to the labelled set. Our proposed `HiLo' networks extract High-level semantic and Low-level domain features, before minimizing the mutual information between the representations. Our intuition is that the clusterings based on domain information and semantic information should be independent. We further extend our method with a specialized domain augmentation tailored for the GCD task, as well as a curriculum learning approach. Finally, we construct a benchmark from corrupted fine-grained datasets as well as a large-scale evaluation on DomainNet with real-world domain shifts, reimplementing a number of GCD baselines in this setting. We demonstrate that HiLo outperforms SoTA category discovery models by a large margin on all evaluations.
翻译:广义类别发现(GCD)是一项具有挑战性的任务,其目标是在给定部分标注数据集的情况下,模型必须对所有未标注实例进行分类,无论这些实例是来自已标注类别还是新类别。在本文中,我们挑战了该任务中一个尚存的假设:所有图像共享同一领域。具体而言,我们引入了一种新任务和方法,以处理当未标注数据也包含与标注集不同领域的图像时的GCD问题。我们提出的`HiLo`网络在最小化表征间互信息之前,分别提取高级语义特征和低级领域特征。我们的直觉是,基于领域信息和语义信息的聚类应当是相互独立的。我们进一步通过一种专为GCD任务设计的领域增强方法以及课程学习策略扩展了我们的方法。最后,我们基于损坏的细粒度数据集构建了一个基准测试,并在DomainNet上进行了大规模的真实世界领域偏移评估,在此设置下重新实现了多个GCD基线模型。实验结果表明,HiLo在所有评估中均以显著优势超越了当前最先进的类别发现模型。