Positive unlabeled learning is a binary classification problem with positive and unlabeled data. It is common in domains where negative labels are costly or impossible to obtain, e.g., medicine and personalized advertising. We apply the locally purified state tensor network to the positive unlabeled learning problem and test our model on the MNIST image and 15 categorical/mixed datasets. On the MNIST dataset, we obtain close to the state-of-the-art results even with very few labeled positive samples. We significantly improve the state-of-the-art on categorical datasets. Further, we show that the agreement fraction between outputs of different models on unlabeled samples is a good indicator of the model's performance. Finally, our method can generate new positive and negative instances, which we demonstrate on simple synthetic datasets.
翻译:正无标签学习是一种仅包含正类数据和无标签数据的二分类问题,常见于难以获取或无法获取负标签的领域(如医学诊断和个性化广告)。我们采用局部纯化态张量网络解决正无标签学习问题,并在MNIST图像数据集及15个分类/混合数据集上测试模型性能。在MNIST数据集上,即便仅使用极少量的标注正样本,我们仍能获得接近最优水平的结果。在分类数据集上,我们显著提升了当前最优方法的性能。进一步研究表明,不同模型在无标签样本上输出的一致性程度是衡量模型性能的良好指标。最后,我们的方法能够生成新的正类和负类样本,并通过简单合成数据集验证了该能力。