A unified approach of Positive and Unlabelled (PU)-learning, Semi-Supervised Learning (SSL), and Open-Set Recognition (OSR) would significantly enhance the development of cost-efficient application-grade classifiers. However, previous attempts have conflated the definitions of \mbox{\textit{observed}} and \mbox{\textit{unobserved}} novel categories. Observed novel categories are defined in PU-learning as those in unlabelled training data and exist due to an incomplete set of category labels for the training set. In contrast, unobserved novel categories are defined in OSR as those that only exist in the testing data and represent new and interesting patterns that emerge over time. To maintain safe and practical classifier development, models must generalise the difference between these novel category types. In this letter, we thoroughly review the relevant machine learning research fields to propose a new unified machine learning policy called Open-set Learning with Augmented Categories by exploiting Unlabelled data or Open-LACU. Specifically, Open-LACU requires models to accurately classify $K > 1$ number of labelled categories while simultaneously detecting and separating observed novel categories into the augmented background category ($K + 1$) and further detecting and separating unobserved novel categories into the augmented unknown category ($K + 2$). Open-LACU is the first machine learning policy to generalise observed and unobserved novel categories. The significance of Open-LACU is also highlighted by discussing its application in semantic segmentation of remote sensing images, object detection within medical radiology images and disease identification through cough sound analysis.
翻译:正例与未标注学习(PU-learning)、半监督学习(SSL)和开放集识别(OSR)的统一方法将显著促进经济高效的应用级分类器的发展。然而,先前的尝试混淆了“观测”和“未观测”新颖类别的定义。观测新颖类别在PU-learning中被定义为未标记训练数据中的类别,其存在是由于训练集的类别标签不完整。相反,未观测新颖类别在OSR中被定义为仅存在于测试数据中的类别,代表了随时间涌现的新颖有趣模式。为确保安全且实用的分类器开发,模型必须泛化这些新颖类别类型之间的差异。本文通过系统梳理相关机器学习研究领域,提出了一种新的统一机器学习策略——基于未标注数据增强类别的开放集学习(Open-LACU)。具体而言,Open-LACU要求模型精确分类$K > 1$个已标注类别,同时检测并分离观测新颖类别至增强背景类别($K + 1$),并进一步检测与分离未观测新颖类别至增强未知类别($K + 2$)。Open-LACU是首个能够泛化观测与未观测新颖类别的机器学习策略。本文通过讨论其在遥感图像语义分割、医学放射图像目标检测以及咳嗽声音分析疾病识别中的应用,进一步凸显了Open-LACU的重要性。