High-dimensional images, known for their rich semantic information, are widely applied in remote sensing and other fields. The spatial information in these images reflects the object's texture features, while the spectral information reveals the potential spectral representations across different bands. Currently, the understanding of high-dimensional images remains limited to a single-domain perspective with performance degradation. Motivated by the masking texture effect observed in the human visual system, we present a multi-domain diffusion-driven feature learning network (MDFL) , a scheme to redefine the effective information domain that the model really focuses on. This method employs diffusion-based posterior sampling to explicitly consider joint information interactions between the high-dimensional manifold structures in the spectral, spatial, and frequency domains, thereby eliminating the influence of masking texture effects in visual models. Additionally, we introduce a feature reuse mechanism to gather deep and raw features of high-dimensional data. We demonstrate that MDFL significantly improves the feature extraction performance of high-dimensional data, thereby providing a powerful aid for revealing the intrinsic patterns and structures of such data. The experimental results on three multi-modal remote sensing datasets show that MDFL reaches an average overall accuracy of 98.25%, outperforming various state-of-the-art baseline schemes. The code will be released, contributing to the computer vision community.
翻译:摘要:高维图像以其丰富的语义信息而闻名,广泛应用于遥感等领域。这些图像中的空间信息反映了物体的纹理特征,而光谱信息则揭示了不同波段间潜在的光谱表征。当前,对高维图像的理解仍局限于单域视角,且存在性能退化问题。受人类视觉系统中掩膜纹理效应的启发,我们提出了一种多域扩散驱动特征学习网络(MDFL),该方案旨在重新定义模型真正关注的有效信息域。该方法采用基于扩散的后验采样,显式考虑光谱、空间和频率域中高维流形结构间的联合信息交互,从而消除视觉模型中掩膜纹理效应的影响。此外,我们引入特征复用机制来收集高维数据的深层与原始特征。实验证明,MDFL显著提升了高维数据的特征提取性能,为揭示此类数据的内在规律与结构提供了有力支撑。在三个多模态遥感数据集上的实验结果表明,MDFL的平均总体精度达到98.25%,优于多种最先进的基线方案。相关代码将公开,以促进计算机视觉领域的发展。