Multimodal imaging and correlative analysis typically require image alignment. Contrastive learning can generate representations of multimodal images, reducing the challenging task of multimodal image registration to a monomodal one. Previously, additional supervision on intermediate layers in contrastive learning has improved biomedical image classification. We evaluate if a similar approach improves representations learned for registration to boost registration performance. We explore three approaches to add contrastive supervision to the latent features of the bottleneck layer in the U-Nets encoding the multimodal images and evaluate three different critic functions. Our results show that representations learned without additional supervision on latent features perform best in the downstream task of registration on two public biomedical datasets. We investigate the performance drop by exploiting recent insights in contrastive learning in classification and self-supervised learning. We visualize the spatial relations of the learned representations by means of multidimensional scaling, and show that additional supervision on the bottleneck layer can lead to partial dimensional collapse of the intermediate embedding space.
翻译:多模态成像及关联分析通常需要图像对齐。对比学习可生成多模态图像的表征,将多模态图像配准这一难题简化为单模态配准任务。此前研究表明,对对比学习中间层施加额外监督可提升生物医学图像分类性能。本研究评估了类似方法能否改善用于配准的表征学习以提升配准性能。我们探索了三种在编码多模态图像的U-Net瓶颈层潜在特征中添加对比监督的方法,并评估了三种不同的评判函数。实验结果表明,在两个公开生物医学数据集上,未对潜在特征施加额外监督所学习的表征在下游配准任务中表现最佳。我们通过借鉴分类领域对比学习与自监督学习的最新研究进展,分析了性能下降的原因。利用多维尺度分析对所学表征的空间关系进行可视化,结果显示对瓶颈层施加额外监督可能导致中间嵌入空间出现部分维度坍缩。