Spatial attention mechanism has been widely incorporated into deep convolutional neural networks (CNNs) via long-range dependency capturing, significantly lifting the performance in computer vision, but it may perform poorly in medical imaging. Unfortunately, existing efforts are often unaware that long-range dependency capturing has limitations in highlighting subtle lesion regions, neglecting to exploit the potential of multi-scale pixel context information to improve the representational capability of CNNs. In this paper, we propose a practical yet lightweight architectural unit, Pyramid Pixel Context Recalibration (PPCR) module, which exploits multi-scale pixel context information to recalibrate pixel position in a pixel-independent manner adaptively. PPCR first designs a cross-channel pyramid pooling to aggregate multi-scale pixel context information, then eliminates the inconsistency among them by the well-designed pixel normalization, and finally estimates per pixel attention weight via a pixel context integration. PPCR can be flexibly plugged into modern CNNs with negligible overhead. Extensive experiments on five medical image datasets and CIFAR benchmarks empirically demonstrate the superiority and generalization of PPCR over state-of-the-art attention methods. The in-depth analyses explain the inherent behavior of PPCR in the decision-making process, improving the interpretability of CNNs.
翻译:空间注意力机制通过长程依赖捕获已被广泛融入深度卷积神经网络(CNNs),显著提升了计算机视觉的性能,但在医学成像中可能表现不佳。遗憾的是,现有工作往往未意识到长程依赖捕获在突出显示微小病变区域方面的局限性,忽视了利用多尺度像素上下文信息来提升CNNs表示能力的潜力。本文提出了一种实用且轻量级的架构单元——金字塔像素上下文重校准(PPCR)模块,该模块利用多尺度像素上下文信息以像素独立的方式自适应地重校准像素位置。PPCR首先设计了一个跨通道金字塔池化来聚合多尺度像素上下文信息,然后通过精心设计的像素归一化消除它们之间的不一致性,最后通过像素上下文集成估算每个像素的注意力权重。PPCR可以灵活地插入现代CNNs中,且开销可忽略不计。在五个医学图像数据集和CIFAR基准上的大量实验实证证明了PPCR相对于最先进注意力方法的优越性和泛化能力。深入分析揭示了PPCR在决策过程中的内在行为,提高了CNNs的可解释性。