Data collected by different modalities can provide a wealth of complementary information, such as hyperspectral image (HSI) to offer rich spectral-spatial properties, synthetic aperture radar (SAR) to provide structural information about the Earth's surface, and light detection and ranging (LiDAR) to cover altitude information about ground elevation. Therefore, a natural idea is to combine multimodal images for refined and accurate land-cover interpretation. Although many efforts have been attempted to achieve multi-source remote sensing image classification, there are still three issues as follows: 1) indiscriminate feature representation without sufficiently considering modal heterogeneity, 2) abundant features and complex computations associated with modeling long-range dependencies, and 3) overfitting phenomenon caused by sparsely labeled samples. To overcome the above barriers, a transformer-based heterogeneously salient graph representation (THSGR) approach is proposed in this paper. First, a multimodal heterogeneous graph encoder is presented to encode distinctively non-Euclidean structural features from heterogeneous data. Then, a self-attention-free multi-convolutional modulator is designed for effective and efficient long-term dependency modeling. Finally, a mean forward is put forward in order to avoid overfitting. Based on the above structures, the proposed model is able to break through modal gaps to obtain differentiated graph representation with competitive time cost, even for a small fraction of training samples. Experiments and analyses on three benchmark datasets with various state-of-the-art (SOTA) methods show the performance of the proposed approach.
翻译:不同模态采集的数据能够提供丰富的互补信息,例如高光谱图像(HSI)提供丰富的光谱-空间特性,合成孔径雷达(SAR)提供地表结构信息,激光雷达(LiDAR)涵盖地面高程的高度信息。因此,结合多模态图像以实现精细准确的土地覆盖解译是一个自然的思路。尽管已有许多研究尝试实现多源遥感图像分类,但仍存在以下三个问题:1)未充分考虑模态异质性导致的特征表示缺乏区分度;2)建模长程依赖时伴随的丰富特征与复杂计算;3)稀疏标注样本引起的过拟合现象。为克服上述障碍,本文提出了一种基于Transformer的异构显著图表示(THSGR)方法。首先,提出了一种多模态异构图编码器,用于从异构数据中编码具有区分度的非欧几里得结构特征。其次,设计了一种无自注意力的多卷积调制器,以实现高效的长程依赖建模。最后,提出了一种均值前向传播策略以避免过拟合。基于以上结构,即使仅使用少量训练样本,所提模型仍能以具有竞争力的时间成本突破模态间隙,获得具有区分度的图表示。在三个基准数据集上与多种前沿(SOTA)方法进行的实验与分析验证了所提方法的性能。