Robots operating in human-centered environments, such as retail stores, restaurants, and households, are often required to distinguish between similar objects in different contexts with a high degree of accuracy. However, fine-grained object recognition remains a challenge in robotics due to the high intra-category and low inter-category dissimilarities. In addition, the limited number of fine-grained 3D datasets poses a significant problem in addressing this issue effectively. In this paper, we propose a hybrid multi-modal Vision Transformer (ViT) and Convolutional Neural Networks (CNN) approach to improve the performance of fine-grained visual classification (FGVC). To address the shortage of FGVC 3D datasets, we generated two synthetic datasets. The first dataset consists of 20 categories related to restaurants with a total of 100 instances, while the second dataset contains 120 shoe instances. Our approach was evaluated on both datasets, and the results indicate that it outperforms both CNN-only and ViT-only baselines, achieving a recognition accuracy of 94.50 % and 93.51 % on the restaurant and shoe datasets, respectively. Additionally, we have made our FGVC RGB-D datasets available to the research community to enable further experimentation and advancement. Furthermore, we successfully integrated our proposed method with a robot framework and demonstrated its potential as a fine-grained perception tool in both simulated and real-world robotic scenarios.
翻译:在人类为中心的环境(如零售商店、餐厅和家庭)中运行的机器人,通常需要高精度地识别不同情境下的相似物体。然而,由于细粒度物体识别中类别内差异大、类别间差异小的特性,该任务在机器人领域仍具挑战性。此外,现有细粒度三维数据集的匮乏严重制约了该问题的有效解决。本文提出一种混合多模态视觉Transformer(ViT)与卷积神经网络(CNN)方法,以提升细粒度视觉分类(FGVC)性能。针对FGVC三维数据集不足的问题,我们生成了两个合成数据集:第一个数据集包含20个与餐厅相关的类别,共100个实例;第二个数据集包含120个鞋子实例。我们在两个数据集上评估了所提方法,结果表明其性能优于仅使用CNN或ViT的基线模型,在餐厅和鞋子数据集上的识别准确率分别达到94.50%和93.51%。此外,我们已将FGVC RGB-D数据集公开,以促进研究社区的进一步实验与进步。同时,我们成功将该方法集成至机器人框架中,并在仿真与真实机器人场景中验证了其作为细粒度感知工具的潜力。