Most commonly used benchmark datasets for computer vision contain irrelevant images, near duplicates, and label errors. Consequently, model performance on these benchmarks may not be an accurate estimate of generalization ability. This is a particularly acute concern in computer vision for medicine where datasets are typically small, stakes are high, and annotation processes are expensive and error-prone. In this paper, we propose SelfClean, a general procedure to clean up image datasets exploiting a latent space learned with self-supervision. By relying on self-supervised learning, our approach focuses on intrinsic properties of the data and avoids annotation biases. We formulate dataset cleaning as either a set of ranking problems, where human experts can make decisions with significantly reduced effort, or a set of scoring problems, where decisions can be fully automated based on score distributions. We compare SelfClean against other algorithms on common computer vision benchmarks enhanced with synthetic noise and demonstrate state-of-the-art performance on detecting irrelevant images, near duplicates, and label errors. In addition, we apply our method to multiple image datasets and confirm an improvement in evaluation reliability.
翻译:大多数常用的计算机视觉基准数据集包含不相关图像、近似重复图像和标签错误。因此,模型在这些基准上的性能可能无法准确估计其泛化能力。这一担忧在医学计算机视觉领域尤为突出,因为该领域数据集通常较小、风险较高且标注过程昂贵且易出错。本文提出SelfClean,一种利用自监督学习潜空间来清理图像数据集的通用流程。通过依赖自监督学习,我们的方法聚焦于数据的内在属性,避免了标注偏差。我们将数据集清洗表述为一系列排序问题(人类专家可在此过程中显著减少工作量)或评分问题(可基于分布完全自动化决策)。我们将SelfClean与其他算法在添加了合成噪声的常见计算机视觉基准上进行对比,并展示了该方法在检测不相关图像、近似重复图像和标签错误方面的最佳性能。此外,我们将该方法应用于多个图像数据集,并证实了评估可靠性的提升。