Learning high-quality, self-supervised, visual representations is essential to advance the role of computer vision in biomedical microscopy and clinical medicine. Previous work has focused on self-supervised representation learning (SSL) methods developed for instance discrimination and applied them directly to image patches, or fields-of-view, sampled from gigapixel whole-slide images (WSIs) used for cancer diagnosis. However, this strategy is limited because it (1) assumes patches from the same patient are independent, (2) neglects the patient-slide-patch hierarchy of clinical biomedical microscopy, and (3) requires strong data augmentations that can degrade downstream performance. Importantly, sampled patches from WSIs of a patient's tumor are a diverse set of image examples that capture the same underlying cancer diagnosis. This motivated HiDisc, a data-driven method that leverages the inherent patient-slide-patch hierarchy of clinical biomedical microscopy to define a hierarchical discriminative learning task that implicitly learns features of the underlying diagnosis. HiDisc uses a self-supervised contrastive learning framework in which positive patch pairs are defined based on a common ancestry in the data hierarchy, and a unified patch, slide, and patient discriminative learning objective is used for visual SSL. We benchmark HiDisc visual representations on two vision tasks using two biomedical microscopy datasets, and demonstrate that (1) HiDisc pretraining outperforms current state-of-the-art self-supervised pretraining methods for cancer diagnosis and genetic mutation prediction, and (2) HiDisc learns high-quality visual representations using natural patch diversity without strong data augmentations.
翻译:高质量自监督视觉表征的学习对于推动计算机视觉在生物医学显微成像和临床医学中的作用至关重要。以往研究主要针对实例判别任务开发自监督表征学习方法(SSL),并将其直接应用于从用于癌症诊断的十亿像素全切片图像(WSI)中采样的图像块(即视野)。然而,这种策略存在局限性,因为其(1)假设同一患者的图像块相互独立,(2)忽略了临床生物医学显微成像中“患者-切片-图像块”的层级结构,以及(3)需要可能降低下游任务性能的强数据增强。值得注意的是,从患者肿瘤WSI中采样的图像块构成了一组多样的图像样本,它们捕捉到同一潜在的癌症诊断结果。基于此,我们提出HiDisc方法——一种数据驱动方法,它利用临床生物医学显微成像固有的“患者-切片-图像块”层级结构来定义层级判别学习任务,从而隐式学习诊断结果的潜在特征。HiDisc采用自监督对比学习框架,其中基于数据层级中的共同祖先定义正样本图像块对,并采用统一的图像块、切片和患者判别学习目标进行视觉SSL。我们在两个生物医学显微成像数据集上,针对两项视觉任务对HiDisc的视觉表征进行基准测试,结果表明:(1)HiDisc预训练方法在癌症诊断和基因突变预测任务上优于当前最先进的自监督预训练方法;(2)HiDisc无需强数据增强,仅利用自然图像块多样性即可学习高质量视觉表征。