This paper introduces an automatic affordance reasoning paradigm tailored to minimal semantic inputs, addressing the critical challenges of classifying and manipulating unseen classes of objects in household settings. Inspired by human cognitive processes, our method integrates generative language models and physics-based simulators to foster analytical thinking and creative imagination of novel affordances. Structured with a tripartite framework consisting of analysis, imagination, and evaluation, our system "analyzes" the requested affordance names into interaction-based definitions, "imagines" the virtual scenarios, and "evaluates" the object affordance. If an object is recognized as possessing the requested affordance, our method also predicts the optimal pose for such functionality, and how a potential user can interact with it. Tuned on only a few synthetic examples across 3 affordance classes, our pipeline achieves a very high success rate on affordance classification and functional pose prediction of 8 classes of novel objects, outperforming learning-based baselines. Validation through real robot manipulating experiments demonstrates the practical applicability of the imagined user interaction, showcasing the system's ability to independently conceptualize unseen affordances and interact with new objects and scenarios in everyday settings.
翻译:本文提出一种面向最小语义输入的自动功能推理范式,旨在解决家庭场景中未见物体类别分类与操作的关键挑战。受人类认知过程启发,本方法融合生成式语言模型与基于物理的仿真器,以促进分析性思维与新型功能的创造性想象。系统采用"分析-想象-评估"三部分框架结构:首先将请求的功能名称"分析"为基于交互的定义,然后"想象"虚拟场景,最后"评估"物体的功能属性。若判定物体具备所请求功能,本方法还可预测实现该功能的最优位姿,以及潜在用户与物体的交互方式。仅通过3个功能类别的少量合成样本进行调优,本流水线在8类新型物体的功能分类与功能位姿预测任务上取得了极高成功率,性能超越基于学习的基线方法。通过真实机器人操作实验验证了想象用户交互的实践可行性,展示了系统独立概念化未知功能并与日常环境中新物体、新场景进行交互的能力。