Robotic grasping in cluttered environments remains a significant challenge due to occlusions and complex object arrangements. We have developed ThinkGrasp, a plug-and-play vision-language grasping system that makes use of GPT-4o's advanced contextual reasoning for heavy clutter environment grasping strategies. ThinkGrasp can effectively identify and generate grasp poses for target objects, even when they are heavily obstructed or nearly invisible, by using goal-oriented language to guide the removal of obstructing objects. This approach progressively uncovers the target object and ultimately grasps it with a few steps and a high success rate. In both simulated and real experiments, ThinkGrasp achieved a high success rate and significantly outperformed state-of-the-art methods in heavily cluttered environments or with diverse unseen objects, demonstrating strong generalization capabilities.
翻译:在杂乱环境中进行机器人抓取仍然是重大挑战,因为存在遮挡和复杂的物体排列。我们开发了ThinkGrasp,这是一个即插即用的视觉-语言抓取系统,它利用GPT-4o的先进语境推理能力来实现高密度杂乱环境中的抓取策略。ThinkGrasp能够有效识别目标物体并生成抓取位姿,即使目标物体被严重遮挡或近乎不可见,也能通过使用目标导向语言引导移除遮挡物体。该方法逐步显露目标物体,最终以少量步骤和高成功率完成抓取。在仿真和真实实验中,ThinkGrasp均取得了高成功率,并且在高度杂乱环境或面对多种未见过的物体时显著优于最先进方法,展现出强大的泛化能力。