Conformal prediction is a statistical framework that generates prediction sets containing ground-truth labels with a desired coverage guarantee. The predicted probabilities produced by machine learning models are generally miscalibrated, leading to large prediction sets in conformal prediction. In this paper, we empirically and theoretically show that disregarding the probabilities' value will mitigate the undesirable effect of miscalibrated probability values. Then, we propose a novel algorithm named $\textit{Sorted Adaptive prediction sets}$ (SAPS), which discards all the probability values except for the maximum softmax probability. The key idea behind SAPS is to minimize the dependence of the non-conformity score on the probability values while retaining the uncertainty information. In this manner, SAPS can produce sets of small size and communicate instance-wise uncertainty. Theoretically, we provide a finite-sample coverage guarantee of SAPS and show that the expected value of set size from SAPS is always smaller than APS. Extensive experiments validate that SAPS not only lessens the prediction sets but also broadly enhances the conditional coverage rate and adaptation of prediction sets.
翻译:共形预测是一种统计框架,能够生成包含真实标签且具有期望覆盖保证的预测集。机器学习模型产生的预测概率通常存在校准偏差,导致共形预测中预测集规模过大。本文通过理论与实验表明,忽略概率值本身可缓解概率值校准不良带来的不利影响。在此基础上,我们提出一种名为$\textit{排序自适应预测集}$(SAPS)的新算法,该算法仅保留最大softmax概率,舍弃所有其他概率值。SAPS的核心思想是在保留不确定性信息的同时,最小化非一致性得分对概率值的依赖。通过这种方式,SAPS能够生成小规模预测集并传递实例级的不确定性。理论方面,我们提供了SAPS在有限样本下的覆盖保证,并证明SAPS的预期预测集规模始终小于APS。大量实验验证表明,SAPS不仅能缩减预测集,还能广泛提升条件覆盖率和预测集的自适应能力。