Unifying semi-supervised learning (SSL) and open-set recognition into a single learning policy would facilitate the development of cost-efficient and application-grade classifiers. However, previous attempts do not clarify the difference between unobserved novel categories (those only seen during testing) and observed novel categories (those present in unlabelled training data). This study introduces Open-Set Learning with Augmented Category by Exploiting Unlabelled Data (Open-LACU), the first policy that generalises between both novel category types. We adapt the state-of-the-art OSR method of Margin Generative Adversarial Networks (Margin-GANs) into several Open-LACU configurations, setting the benchmarks for Open-LACU and offering unique insights into novelty detection using Margin-GANs. Finally, we highlight the significance of the Open-LACU policy by discussing the applications of semantic segmentation in remote sensing, object detection in radiology and disease identification through cough analysis. These applications include observed and unobserved novel categories, making Open-LACU essential for training classifiers in these big data domains.
翻译:将半监督学习(SSL)与开放集识别统一到单一学习策略中,有助于开发经济高效且实用级别的分类器。然而,以往的研究未能明确区分未观测到的新颖类别(仅在测试阶段出现)与已观测到的新颖类别(存在于未标注训练数据中)。本研究提出了一种通过利用未标注数据增强类别的开放集学习(Open-LACU)方法,这是首个能够泛化处理两种新颖类别类型的学习策略。我们将先进的开放集识别方法——边际生成对抗网络(Margin-GANs)——适配至多种Open-LACU配置中,建立了Open-LACU的基准测试,并提供了利用边际生成对抗网络进行新颖性检测的独特见解。最后,我们通过讨论语义分割在遥感、目标检测在放射学以及咳嗽分析在疾病识别中的应用,强调了Open-LACU策略的重要性。这些应用涵盖了已观测与未观测的新颖类别,使得Open-LACU在大数据领域中训练分类器时成为关键工具。