Non-contact material identification enables adaptive interaction for embodied intelligence yet faces challenges from geometry-induced variations (e.g., orientation, shape, distance) and single-modality ambiguities. In this paper, we present GaMi, a multimodal material identification system integrating mmWave and acoustic sensing to robustly operate under unconstrained geometric conditions. By leveraging the insight of shared geometric consistency between co-located bimodal sensors, GaMi employs an intra-sample cross-modal subtractive disentanglement framework. By semantically aligning modalities and subtracting the shared geometric context, it isolates intrinsic material features. Furthermore, GaMi incorporates inter-sample contrastive learning to correct the residual interference caused by cross-modal misalignment. Additionally, a pairing-based adaptation strategy between two modalities enables few-shot generalization across devices. Extensive evaluations on 20 materials show that GaMi achieves 95.2% accuracy, outperforming single-modality baselines across unseen geometric conditions.
翻译:非接触式材料识别能够实现具身智能的自适应交互,但面临几何因素(如朝向、形状、距离)引起的纹理变化以及单模态歧义性的挑战。本文提出GaMi,一种融合毫米波与声学传感的多模态材料识别系统,可在无约束几何条件下鲁棒运行。基于同位置双模态传感器共享几何一致性的洞察,GaMi采用样本内跨模态减法解耦框架:通过对齐模态语义并减去共享的几何上下文,分离出固有材料特征。进一步,GaMi引入样本间对比学习以校正跨模态错位导致的残差干扰。此外,双模态间的配对式自适应策略支持设备间的小样本泛化。对20种材料的评估表明,GaMi在未见过的几何条件下实现了95.2%的准确率,优于单模态基线。