Precision medicine fundamentally aims to establish causality between dysregulated biochemical mechanisms and cancer subtypes. Omics-based cancer subtyping has emerged as a revolutionary approach, as different level of omics records the biochemical products of multistep processes in cancers. This paper focuses on fully exploiting the potential of multi-omics data to improve cancer subtyping outcomes, and hence developed MoCLIM, a representation learning framework. MoCLIM independently extracts the informative features from distinct omics modalities. Using a unified representation informed by contrastive learning of different omics modalities, we can well-cluster the subtypes, given cancer, into a lower latent space. This contrast can be interpreted as a projection of inter-omics inference observed in biological networks. Experimental results on six cancer datasets demonstrate that our approach significantly improves data fit and subtyping performance in fewer high-dimensional cancer instances. Moreover, our framework incorporates various medical evaluations as the final component, providing high interpretability in medical analysis.
翻译:精准医学的核心在于建立失调生化机制与癌症亚型之间的因果关系。基于组学的癌症亚型分类已成为革命性方法,不同层次的组学数据记录了癌症多步过程中的生化产物。本文聚焦于充分挖掘多组学数据的潜力以提升癌症亚型分类效果,由此提出了表征学习框架MoCLIM。该框架从不同组学模态中独立提取信息特征,通过对比学习融合各模态的统一表征,可在低维隐空间中对特定癌症亚型进行有效聚类。这种对比可被解释为生物网络中组间推理的投影映射。在六个癌症数据集上的实验结果表明,本方法在稀疏高维癌症样本中显著提升了数据拟合度与亚型分类性能。此外,框架整合了多元医学评估作为最终模块,为医学分析提供了高可解释性。