As a crucial step toward real-world learning scenarios with changing environments, dataset shift theory and invariant representation learning algorithm have been extensively studied to relax the identical distribution assumption in classical learning setting. Among the different assumptions on the essential of shifting distributions, generalized label shift (GLS) is the latest developed one which shows great potential to deal with the complex factors within the shift. In this paper, we aim to explore the limitations of current dataset shift theory and algorithm, and further provide new insights by presenting a comprehensive understanding of GLS. From theoretical aspect, two informative generalization bounds are derived, and the GLS learner is proved to be sufficiently close to optimal target model from the Bayesian perspective. The main results show the insufficiency of invariant representation learning, and prove the sufficiency and necessity of GLS correction for generalization, which provide theoretical supports and innovations for exploring generalizable model under dataset shift. From methodological aspect, we provide a unified view of existing shift correction frameworks, and propose a kernel embedding-based correction algorithm (KECA) to minimize the generalization error and achieve successful knowledge transfer. Both theoretical results and extensive experiment evaluations demonstrate the sufficiency and necessity of GLS correction for addressing dataset shift and the superiority of proposed algorithm.
翻译:作为迈向环境变化的现实学习场景的关键一步,数据集偏移理论与不变表示学习算法已被广泛研究,以放宽经典学习设置中的同分布假设。在关于偏移分布本质的不同假设中,广义标签偏移(GLS)是最新发展的理论框架,展现出处理偏移中复杂因素的巨大潜力。本文旨在探索当前数据集偏移理论与算法的局限性,并通过呈现对GLS的全面理解提供新的见解。从理论层面,我们推导出两个具有信息量的泛化界,并从贝叶斯视角证明GLS学习器充分接近最优目标模型。主要结果表明了不变表示学习的不足性,并证明了GLS校正对泛化的充分必要性,这为探索数据集偏移下可泛化模型提供了理论支撑与创新思路。从方法层面,我们提供了现有偏移校正框架的统一视角,并提出一种基于核嵌入的校正算法(KECA)以最小化泛化误差并实现成功的知识迁移。理论结果与大量实验评估共同证明了GLS校正对于解决数据集偏移的充分必要性以及所提算法的优越性。