For classifying digital whole slide images in the absence of pixel level annotation, typically multiple instance learning methods are applied. Due to the generic applicability, such methods are currently of very high interest in the research community, however, the issue of data augmentation in this context is rarely explored. Here we investigate linear and multilinear interpolation between feature vectors, a data augmentation technique, which proved to be capable of improving the generalization performance classification networks and also for multiple instance learning. Experiments, however, have been performed on only two rather small data sets and one specific feature extraction approach so far and a strong dependence on the data set has been identified. Here we conduct a large study incorporating 10 different data set configurations, two different feature extraction approaches (supervised and self-supervised), stain normalization and two multiple instance learning architectures. The results showed an extraordinarily high variability in the effect of the method. We identified several interesting aspects to bring light into the darkness and identified novel promising fields of research.
翻译:在缺乏像素级标注的数字全切片图像分类任务中,通常采用多实例学习方法。由于该方法具有通用适用性,目前已成为研究界高度关注的课题,但在此背景下数据增强问题却鲜有探讨。本文研究了一种基于特征向量线性与多线性插值的数据增强技术,该技术已被证明能够提升分类网络的泛化性能,同时适用于多实例学习场景。然而,现有实验仅基于两个较小数据集和一种特定特征提取方法展开,且发现其结果对数据集存在强烈依赖性。本研究开展了大规模实验,涵盖10种不同数据集配置、两种特征提取方法(监督与自监督学习)、染色归一化处理及两种多实例学习架构。实验结果显示该方法的效果呈现极高变异性。我们揭示了若干值得关注的现象以阐明该领域的研究盲点,并发现了具有前景的新型研究方向。