Instance segmentation in electron microscopy (EM) volumes is tough due to complex shapes and sparse annotations. Self-supervised learning helps but still struggles with intricate visual patterns in EM. To address this, we propose a pretraining framework that enhances multiscale consistency in EM volumes. Our approach leverages a Siamese network architecture, integrating both strong and weak data augmentations to effectively extract multiscale features. We uphold voxel-level coherence by reconstructing the original input data from these augmented instances. Furthermore, we incorporate cross-attention mechanisms to facilitate fine-grained feature alignment between these augmentations. Finally, we apply contrastive learning techniques across a feature pyramid, allowing us to distill distinctive representations spanning various scales. After pretraining on four large-scale EM datasets, our framework significantly improves downstream tasks like neuron and mitochondria segmentation, especially with limited finetuning data. It effectively captures voxel and feature consistency, showing promise for learning transferable representations for EM analysis.
翻译:电子显微镜(EM)体数据中的实例分割因复杂形状和稀疏标注而极具挑战。自监督学习虽有所助益,但在处理EM中复杂的视觉模式时仍存在困难。为此,我们提出一种预训练框架,旨在增强EM体数据中的多尺度一致性。该方法采用孪生网络架构,整合强数据增强与弱数据增强,以高效提取多尺度特征。我们通过对这些增强后的实例重建原始输入数据,维持体素级的一致性。此外,引入交叉注意力机制以促进增强之间的细粒度特征对齐。最后,在特征金字塔上应用对比学习技术,从而提取跨越不同尺度的判别性表征。在四个大规模EM数据集上预训练后,我们的框架显著提升了下游任务(如神经元和线粒体分割)的性能,尤其在微调数据有限的情况下,能够有效捕捉体素与特征的一致性,展现出为EM分析学习可迁移表征的潜力。