Enabling robots to understand and execute tasks from natural language commands while maintaining data efficiency remains challenging. Foundation models such as vision-language-action (VLA) and vision-language models (VLMs) provide intuitive interaction channels but require extensive data; task-parameterized imitation learning achieves data efficiency but lacks natural language grounding. This work bridges this gap through a modular architecture combining task-parameterized kernelized movement primitives (TP-KMPs) with pretrained VLMs. During learning, skills are acquired from 2 to 5 kinesthetic demonstrations, and the VLM generates skill schemas describing each skill's parameters and preconditions. During execution, the VLM interprets commands to select skills, reason about parameter bindings, and create novel behaviors through covariance-weighted composition. When no skill or composition suffices, the system identifies capability gaps and requests targeted demonstrations, all without fine-tuning. Validation on a 7-DoF manipulator shows success rates of 73.3%-100% in scenarios requiring skill selection, composition, and active learning.
翻译:摘要:使机器人能够理解自然语言指令并执行相应任务,同时保持数据效率仍是一项挑战。视觉-语言-动作(VLA)和视觉-语言模型(VLM)等基础模型提供了直观的交互通道,但需要大量数据;任务参数化模仿学习实现了数据效率,却缺乏自然语言基础。本研究通过一种模块化架构弥合了这一差距,该架构将任务参数化核化运动基元(TP-KMP)与预训练VLM相结合。在学习阶段,技能通过2至5次动觉示教获得,VLM生成描述每个技能参数和前提条件的技能图式。在执行阶段,VLM解析指令以选择技能、推理参数绑定,并通过协方差加权组合创造新颖行为。当无技能或组合可满足需求时,系统可识别能力缺口并请求针对性示教,整个过程无需微调。在7自由度机械臂上的验证显示,在需要技能选择、组合和主动学习的场景中,成功率可达73.3%-100%。