Positive-unlabeled learning (PUL) aims at learning a binary classifier from only positive and unlabeled training data. Even though real-world applications often involve imbalanced datasets where the majority of examples belong to one class, most contemporary approaches to PUL do not investigate performance in this setting, thus severely limiting their applicability in practice. In this work, we thus propose to tackle the issues of imbalanced datasets and model calibration in a PUL setting through an uncertainty-aware pseudo-labeling procedure (PUUPL): by boosting the signal from the minority class, pseudo-labeling expands the labeled dataset with new samples from the unlabeled set, while explicit uncertainty quantification prevents the emergence of harmful confirmation bias leading to increased predictive performance. Within a series of experiments, PUUPL yields substantial performance gains in highly imbalanced settings while also showing strong performance in balanced PU scenarios across recent baselines. We furthermore provide ablations and sensitivity analyses to shed light on PUUPL's several ingredients. Finally, a real-world application with an imbalanced dataset confirms the advantage of our approach.
翻译:正-无标签学习(PUL)旨在仅从正样本和无标签训练数据中学习二分类器。尽管现实应用常面临多数样本属于某一类别的类别不平衡数据集,但当前大多数PUL方法并未研究该场景下的性能表现,从而严重限制了其实际应用。为此,本文通过一种不确定性感知的伪标签标注过程(PUUPL)来解决PUL场景中的类别不平衡与模型校准问题:通过增强少数类信号,伪标签利用无标签集中的新样本扩展带标签数据集,同时显式的不确定性量化可防止有害确认偏差的产生,从而提升预测性能。在系列实验中,PUUPL在高度不平衡场景中取得了显著性能提升,同时在平衡PUL场景中较近期基线方法也展现出强劲表现。我们进一步通过消融实验与敏感性分析揭示PUUPL各组成部分的作用机理。最后,基于不平衡数据集的实际应用验证了本方法的优势。