Extreme classification (XC) involves predicting over large numbers of classes (thousands to millions), with real-world applications like news article classification and e-commerce product tagging. The zero-shot version of this task requires generalization to novel classes without additional supervision. In this paper, we develop SemSup-XC, a model that achieves state-of-the-art zero-shot and few-shot performance on three XC datasets derived from legal, e-commerce, and Wikipedia data. To develop SemSup-XC, we use automatically collected semantic class descriptions to represent classes and facilitate generalization through a novel hybrid matching module that matches input instances to class descriptions using a combination of semantic and lexical similarity. Trained with contrastive learning, SemSup-XC significantly outperforms baselines and establishes state-of-the-art performance on all three datasets considered, gaining up to 12 precision points on zero-shot and more than 10 precision points on one-shot tests, with similar gains for recall@10. Our ablation studies highlight the relative importance of our hybrid matching module and automatically collected class descriptions.
翻译:摘要:极端分类(Extreme Classification, XC)涉及对大量类别(数千至数百万个)进行预测,其实际应用包括新闻文章分类和电子商务产品标签。该任务的零样本版本要求模型能够在无需额外监督的情况下泛化至新类别。本文提出了SemSup-XC模型,该模型在法律、电子商务及维基百科数据衍生的三个XC数据集上取得了零样本与小样本场景下的最优性能。为构建SemSup-XC,我们利用自动收集的语义类别描述表示类别,并通过一种新颖的混合匹配模块——该模块结合语义相似性与词法相似性将输入实例与类别描述进行匹配——增强泛化能力。经对比学习训练后,SemSup-XC显著超越基线方法,并在所有三个数据集上确立了最优性能:零样本测试中精度提升高达12个百分点,单样本测试中精度提升超过10个百分点,且召回率@10指标亦有类似提升。我们的消融研究揭示了混合匹配模块与自动收集的类别描述各自的相对重要性。