Task-oriented grasping (TOG) is more challenging than simple object grasping because it requires precise identification of object parts and careful selection of grasping areas to ensure effective and robust manipulation. While recent approaches have trained large-scale vision-language models to integrate part-level object segmentation with task-aware grasp planning, their instability in part recognition and grasp inference limits their ability to generalize across diverse objects and tasks. To address this issue, we introduce a novel, geometry-centric strategy for more generalizable TOG that does not rely on semantic features from visual recognition, effectively overcoming the viewpoint sensitivity of model-based approaches. Our main proposals include: 1) an object-part-task ontology for functional part selection based on intuitive human commands, constructed using a Large Language Model (LLM); 2) a sampling-based geometric analysis method for identifying the selected object part from observed point clouds, incorporating multiple point distribution and distance metrics; and 3) a similarity matching framework for imitative grasp planning, utilizing similar known objects with pre-existing segmentation and grasping knowledge as references to guide the planning for unknown targets. We validate the high accuracy of our approach in functional part selection, identification, and grasp generation through real-world experiments. Additionally, we demonstrate the method's generalization capabilities to novel-category objects by extending existing ontological knowledge, showcasing its adaptability to a broad range of objects and tasks.
翻译:任务导向抓取(TOG)比简单目标抓取更具挑战性,因为它需要精确识别目标部件并谨慎选择抓取区域,以确保高效且鲁棒的操作。尽管近期研究通过训练大规模视觉-语言模型来实现部件级目标分割与任务感知抓取规划的整合,但其在部件识别与抓取推理中的不稳定性限制了跨多样目标及任务的泛化能力。为解决此问题,我们提出一种以几何为中心的新颖策略,无需依赖视觉识别的语义特征即可实现更泛化的TOG,有效克服了基于模型方法对视角的敏感性。主要贡献包括:1)基于大型语言模型(LLM)构建目标-部件-任务本体,实现根据直观人类指令的功能性部件选择;2)提出采样驱动的几何分析方法,融合多类点分布与距离度量,从观测点云中识别选定目标部件;3)构建模仿抓取规划的相似性匹配框架,利用具有预分割与抓取知识的已知相似目标作为参考,指导未知目标的规划。通过真实世界实验验证了本方法在功能性部件选择、识别及抓取生成方面的高精度。此外,通过扩展现有本体知识,证明了该方法对新型类别目标的泛化能力,展示了其在广泛目标与任务中的适应性。