Technological advances in medical data collection such as high-resolution histopathology and high-throughput genomic sequencing have contributed to the rising requirement for multi-modal biomedical modelling, specifically for image, tabular, and graph data. Most multi-modal deep learning approaches use modality-specific architectures that are trained separately and cannot capture the crucial cross-modal information that motivates the integration of different data sources. This paper presents the Hybrid Early-fusion Attention Learning Network (HEALNet): a flexible multi-modal fusion architecture, which a) preserves modality-specific structural information, b) captures the cross-modal interactions and structural information in a shared latent space, c) can effectively handle missing modalities during training and inference, and d) enables intuitive model inspection by learning on the raw data input instead of opaque embeddings. We conduct multi-modal survival analysis on Whole Slide Images and Multi-omic data on four cancer cohorts of The Cancer Genome Atlas (TCGA). HEALNet achieves state-of-the-art performance, substantially improving over both uni-modal and recent multi-modal baselines, whilst being robust in scenarios with missing modalities.
翻译:医学数据采集技术的进步,如高分辨率组织病理学和高通量基因组测序,推动了多模态生物医学建模需求的日益增长,特别是针对图像、表格和图数据。大多数多模态深度学习方法采用特定模态的架构,这些架构单独训练且无法捕捉推动不同数据源整合的关键跨模态信息。本文提出混合早期融合注意力学习网络(HEALNet):一种灵活的多模态融合架构,它能够a)保留模态特定的结构信息,b)在共享潜在空间中捕捉跨模态交互与结构信息,c)在训练和推理过程中有效处理缺失模态,d)通过对原始数据输入而非不透明嵌入进行学习,实现直观的模型检查。我们在癌症基因组图谱(TCGA)的四个癌症队列上,对全切片图像和多组学数据进行多模态生存分析。HEALNet实现了最先进的性能,显著优于单模态和近期多模态基线方法,同时在模态缺失场景中保持鲁棒性。