Unsupervised domain adaptation addresses the problem of classifying data in an unlabeled target domain, given labeled source domain data that share a common label space but follow a different distribution. Most of the recent methods take the approach of explicitly aligning feature distributions between the two domains. Differently, motivated by the fundamental assumption for domain adaptability, we re-cast the domain adaptation problem as discriminative clustering of target data, given strong privileged information provided by the closely related, labeled source data. Technically, we use clustering objectives based on a robust variant of entropy minimization that adaptively filters target data, a soft Fisher-like criterion, and additionally the cluster ordering via centroid classification. To distill discriminative source information for target clustering, we propose to jointly train the network using parallel, supervised learning objectives over labeled source data. We term our method of distilled discriminative clustering for domain adaptation as DisClusterDA. We also give geometric intuition that illustrates how constituent objectives of DisClusterDA help learn class-wisely pure, compact feature distributions. We conduct careful ablation studies and extensive experiments on five popular benchmark datasets, including a multi-source domain adaptation one. Based on commonly used backbone networks, DisClusterDA outperforms existing methods on these benchmarks. It is also interesting to observe that in our DisClusterDA framework, adding an additional loss term that explicitly learns to align class-level feature distributions across domains does harm to the adaptation performance, though more careful studies in different algorithmic frameworks are to be conducted.
翻译:无监督域适应解决的是在给定共享相同标签空间但分布不同的有标签源域数据情况下,对无标签目标域数据进行分类的问题。近期大多数方法采用显式对齐两个域之间特征分布的策略。与之不同,受域适应基本假设的启发,我们将域适应问题重新表述为对目标数据的判别聚类,其中利用了密切相关且有标签的源域数据提供的强特权信息。技术上,我们采用基于熵最小化鲁棒变体的聚类目标,该变体能自适应过滤目标数据,并结合软Fisher准则以及通过质心分类实现的聚类排序。为蒸馏面向目标聚类的判别性源域信息,我们提出在并行监督学习目标下对有标签源域数据联合训练网络。我们将这种用于域适应的蒸馏判别聚类方法命名为DisClusterDA。我们还提供了几何直观解释,阐明DisClusterDA的各组成部分目标如何帮助学习类别纯净且紧凑的特征分布。我们在五个流行的基准数据集(包括一个多源域适应数据集)上进行了细致的消融研究和广泛实验。基于常用骨干网络,DisClusterDA在这些基准上优于现有方法。有趣的是,在DisClusterDA框架中,额外添加显式学习对齐跨域类别级特征分布的损失项反而会损害适应性能,尽管这需要在不同算法框架下进行更细致的研究。