Traditional text classification typically categorizes texts into pre-defined coarse-grained classes, from which the produced models cannot handle the real-world scenario where finer categories emerge periodically for accurate services. In this work, we investigate the setting where fine-grained classification is done only using the annotation of coarse-grained categories and the coarse-to-fine mapping. We propose a lightweight contrastive clustering-based bootstrapping method to iteratively refine the labels of passages. During clustering, it pulls away negative passage-prototype pairs under the guidance of the mapping from both global and local perspectives. Experiments on NYT and 20News show that our method outperforms the state-of-the-art methods by a large margin.
翻译:传统文本分类通常将文本划分为预定义的粗粒度类别,由此产生的模型无法应对实际场景中需要精准服务时周期性出现的新细粒度类别。本研究探究仅利用粗粒度类别标注及粗到细映射关系实现细粒度分类的设定。我们提出一种基于轻量级对比聚类的引导方法,通过迭代方式精炼文本段落的标签。在聚类过程中,该方法在全局与局部映射的双重引导下,将负例段落-原型对分离开来。在NYT和20News数据集上的实验表明,本方法以显著优势超越现有最优方法。