Class distribution shifts are particularly challenging for zero-shot classifiers, which rely on representations learned from training classes but are deployed on new, unseen ones. Common causes for such shifts are changes in attributes associated with classes, such as race or gender in person identification. In this work, we propose and analyze a model that adopts this setting, assuming that the attribute responsible for the shift is unknown during training. To address the challenge of learning data representations robust to such shifts, we introduce a framework based on hierarchical sampling to construct synthetic data environments. Despite key differences between the settings, this framework allows us to formulate class distribution shifts in zero-shot learning as out-of-distribution problems. Consequently, we present an algorithm for learning robust representations, and show that our approach significantly improves generalization to diverse class distributions in both simulations and real-world datasets.
翻译:类别分布偏移对零样本分类器构成了特别严峻的挑战,这类分类器依赖于从训练类别中学到的表示,但却应用于新的、未见过的类别。导致此类偏移的常见原因包括与类别相关的属性变化,例如人物识别中的种族或性别属性。在本研究中,我们提出并分析了一个适应这种场景的模型,该模型假设导致偏移的属性在训练期间是未知的。为了应对学习对这类偏移鲁棒的数据表示这一挑战,我们引入了一个基于分层采样的框架来构建合成数据环境。尽管不同场景之间存在关键差异,该框架使我们能够将零样本学习中的类别分布偏移形式化为分布外问题。因此,我们提出了一种学习鲁棒表示的算法,并表明我们的方法在模拟实验和真实数据集中均显著提升了对多样化类别分布的泛化能力。