Selecting informative data points for expert feedback can significantly improve the performance of anomaly detection (AD) in various contexts, such as medical diagnostics or fraud detection. In this paper, we determine a set of theoretical conditions under which anomaly scores generalize from labeled queries to unlabeled data. Motivated by these results, we propose a data labeling strategy with optimal data coverage under labeling budget constraints. In addition, we propose a new learning framework for semi-supervised AD. Extensive experiments on image, tabular, and video data sets show that our approach results in state-of-the-art semi-supervised AD performance under labeling budget constraints.
翻译:选择用于专家反馈的信息性数据点,可在医学诊断或欺诈检测等场景中显著提升异常检测(AD)的性能。本文确定了异常分数从标注查询泛化至未标注数据时所需的理论条件集合。基于这些结论,我们提出了一种在标注预算约束下实现最优数据覆盖率的标注策略,并构建了面向半监督异常检测的新型学习框架。在图像、表格及视频数据集上的大量实验表明,本方法在标注预算约束下达到了半监督异常检测领域最先进的性能水平。