We extend our previous work on Inductive Conformal Prediction (ICP) for multi-label text classification and present a novel approach for addressing the computational inefficiency of the Label Powerset (LP) ICP, arrising when dealing with a high number of unique labels. We present experimental results using the original and the proposed efficient LP-ICP on two English and one Czech language data-sets. Specifically, we apply the LP-ICP on three deep Artificial Neural Network (ANN) classifiers of two types: one based on contextualised (bert) and two on non-contextualised (word2vec) word-embeddings. In the LP-ICP setting we assign nonconformity scores to label-sets from which the corresponding p-values and prediction-sets are determined. Our approach deals with the increased computational burden of LP by eliminating from consideration a significant number of label-sets that will surely have p-values below the specified significance level. This reduces dramatically the computational complexity of the approach while fully respecting the standard CP guarantees. Our experimental results show that the contextualised-based classifier surpasses the non-contextualised-based ones and obtains state-of-the-art performance for all data-sets examined. The good performance of the underlying classifiers is carried on to their ICP counterparts without any significant accuracy loss, but with the added benefits of ICP, i.e. the confidence information encapsulated in the prediction sets. We experimentally demonstrate that the resulting prediction sets can be tight enough to be practically useful even though the set of all possible label-sets contains more than $1e+16$ combinations. Additionally, the empirical error rates of the obtained prediction-sets confirm that our outputs are well-calibrated.
翻译:我们将先前关于归纳共形预测(ICP)在多标签文本分类中的研究工作进行了扩展,并提出了一种新方法,用于解决标签幂集(LP)ICP在处理大量唯一标签时出现的计算低效问题。我们基于原始及所提出的高效LP-ICP方法,在两个英语数据集和一个捷克语数据集上展示了实验结果。具体而言,我们将LP-ICP应用于三种深度人工神经网络(ANN)分类器:一种基于上下文化的词嵌入(BERT),另两种基于非上下文化的词嵌入(Word2Vec)。在LP-ICP设定中,我们为标签集分配非一致性分数,并据此确定相应的p值和预测集。我们的方法通过剔除那些p值必然低于指定显著性水平的大量标签集,有效应对了LP的计算负担。这大幅降低了该方法的计算复杂度,同时完全遵循标准共形预测(CP)的保障。实验结果表明,基于上下文化嵌入的分类器超越非上下文化分类器,并在所有测试数据集上取得了当前最优性能。基础分类器的优异性能被其ICP对应版本继承,且未出现显著的精度损失,但获得了ICP的附加优势,即预测集中蕴含的置信度信息。我们通过实验证明,即使所有可能标签集的组合数超过$1\times10^{16}$,所得预测集仍可足够紧凑以具备实际应用价值。此外,预测集的经验误差率进一步验证了输出结果的良好校准性。