A holistic understanding of object properties across diverse sensory modalities (e.g., visual, audio, and haptic) is essential for tasks ranging from object categorization to complex manipulation. Drawing inspiration from cognitive science studies that emphasize the significance of multi-sensory integration in human perception, we introduce MOSAIC (Multi-modal Object property learning with Self-Attention and Integrated Comprehension), a novel framework designed to facilitate the learning of unified multi-sensory object property representations. While it is undeniable that visual information plays a prominent role, we acknowledge that many fundamental object properties extend beyond the visual domain to encompass attributes like texture, mass distribution, or sounds, which significantly influence how we interact with objects. In MOSAIC, we leverage this profound insight by distilling knowledge from the extensive pre-trained Contrastive Language-Image Pre-training (CLIP) model, aligning these representations not only across vision but also haptic and auditory sensory modalities. Through extensive experiments on a dataset where a humanoid robot interacts with 100 objects across 10 exploratory behaviors, we demonstrate the versatility of MOSAIC in two task families: object categorization and object-fetching tasks. Our results underscore the efficacy of MOSAIC's unified representations, showing competitive performance in category recognition through a simple linear probe setup and excelling in the fetch object task under zero-shot transfer conditions. This work pioneers the application of CLIP-based sensory grounding in robotics, promising a significant leap in multi-sensory perception capabilities for autonomous systems. We have released the code, datasets, and additional results: https://github.com/gtatiya/MOSAIC.
翻译:对物体属性在多种感官模态(例如视觉、听觉和触觉)上的整体性理解,对于从物体分类到复杂操作等任务至关重要。受认知科学中强调多感官整合对人类感知重要性的研究启发,我们引入了MOSAIC(基于自注意力与综合理解的多模态物体属性学习)——一种旨在促进统一多感官物体属性表示学习的新型框架。尽管视觉信息无疑扮演着突出角色,但我们认识到许多基本物体属性超越视觉范畴,涵盖了纹理、质量分布或声音等特征,这些属性显著影响着我们与物体的交互方式。在MOSAIC中,我们利用这一深刻洞见,从大规模预训练的对比语言-图像预训练(CLIP)模型中提炼知识,将这些表示不仅跨视觉模态对齐,还扩展至触觉和听觉感官模态。通过在一个人形机器人以10种探索行为交互100个物体的数据集上进行广泛实验,我们展示了MOSAIC在物体分类与物体抓取这两类任务中的多功能性。我们的结果强调了MOSAIC统一表示的有效性:在简单线性探测设置下,其类别识别性能具有竞争力,并在零样本迁移条件下的抓取物体任务中表现出色。本工作开创性地将基于CLIP的感官基础应用于机器人学,有望显著提升自主系统的多感官感知能力。我们已发布代码、数据集及附加结果:https://github.com/gtatiya/MOSAIC。