In this paper, we consider the intersection of two problems in machine learning: Multi-Source Domain Adaptation (MSDA) and Dataset Distillation (DD). On the one hand, the first considers adapting multiple heterogeneous labeled source domains to an unlabeled target domain. On the other hand, the second attacks the problem of synthesizing a small summary containing all the information about the datasets. We thus consider a new problem called MSDA-DD. To solve it, we adapt previous works in the MSDA literature, such as Wasserstein Barycenter Transport and Dataset Dictionary Learning, as well as DD method Distribution Matching. We thoroughly experiment with this novel problem on four benchmarks (Caltech-Office 10, Tennessee-Eastman Process, Continuous Stirred Tank Reactor, and Case Western Reserve University), where we show that, even with as little as 1 sample per class, one achieves state-of-the-art adaptation performance.
翻译:本文探讨了机器学习中两个问题的交叉:多源域适应(MSDA)与数据集蒸馏(DD)。前者关注将多个异质标注源域自适应地迁移至无标注目标域;后者则致力于合成包含数据集全部信息的小型摘要。为此,我们提出一种名为MSDA-DD的新问题。为解决该问题,我们改进MSDA文献中的先前工作(如Wasserstein重心传输与数据集字典学习)以及DD方法中的分布匹配技术。我们在四个基准测试(Caltech-Office 10、田纳西-伊斯曼过程、连续搅拌釜式反应器及凯斯西储大学数据集)上对该新问题进行了充分实验,结果表明,即使每类仅采用1个样本,仍能达到最先进的适应性能。