This study investigates how text-driven object affordance, which provides prior knowledge about grasp types for each object, affects image-based grasp-type recognition in robot teaching. The researchers created labeled datasets of first-person hand images to examine the impact of object affordance on recognition performance. They evaluated scenarios with real and illusory objects, considering mixed reality teaching conditions where visual object information may be limited. The results demonstrate that object affordance improves image-based recognition by filtering out unlikely grasp types and emphasizing likely ones. The effectiveness of object affordance was more pronounced when there was a stronger bias towards specific grasp types for each object. These findings highlight the significance of object affordance in multimodal robot teaching, regardless of whether real objects are present in the images. Sample code is available on https://github.com/microsoft/arr-grasp-type-recognition.
翻译:本研究探讨了文本驱动的物体可供性(即每种物体抓取类型的先验知识)如何影响机器人示教中基于图像的抓取类型识别。研究人员创建了第一人称手部图像标注数据集,以检验物体可供性对识别性能的影响。他们评估了真实物体与虚拟物体场景,并考虑了混合现实示教条件下视觉物体信息可能受限的情况。结果表明,物体可供性通过滤除不可能的抓取类型并突出可能的抓取类型,提升了基于图像的识别效果。当每种物体对特定抓取类型的偏向性更强时,物体可供性的有效性更为显著。这些发现强调了物体可供性在多模态机器人示教中的重要性,无论图像中是否存在真实物体。示例代码可在 https://github.com/microsoft/arr-grasp-type-recognition 获取。