Jointly harnessing complementary features of multi-modal input data in a common latent space has been found to be beneficial long ago. However, the influence of each modality on the models decision remains a puzzle. This study proposes a deep learning framework for the modality-level interpretation of multimodal earth observation data in an end-to-end fashion. While leveraging an explainable machine learning method, namely Occlusion Sensitivity, the proposed framework investigates the influence of modalities under an early-fusion scenario in which the modalities are fused before the learning process. We show that the task of wilderness mapping largely benefits from auxiliary data such as land cover and night time light data.
翻译:长期以来,人们发现将多模态输入数据的互补特征联合利用至共同潜在空间是有益的。然而,每种模态对模型决策的影响仍是一个未解之谜。本研究提出一种端到端的深度学习框架,用于多模态地球观测数据的模态级解释。该框架利用可解释机器学习方法(即遮挡敏感性),探讨了在早期融合场景下(即各模态在学习过程之前进行融合)模态的影响。我们证明,荒野制图任务在很大程度上受益于辅助数据,例如土地覆盖和夜间灯光数据。