When limited by their own morphologies, humans and some species of animals have the remarkable ability to use objects from the environment toward accomplishing otherwise impossible tasks. Robots might similarly unlock a range of additional capabilities through tool use. Recent techniques for jointly optimizing morphology and control via deep learning are effective at designing locomotion agents. But while outputting a single morphology makes sense for locomotion, manipulation involves a variety of strategies depending on the task goals at hand. A manipulation agent must be capable of rapidly prototyping specialized tools for different goals. Therefore, we propose learning a designer policy, rather than a single design. A designer policy is conditioned on task information and outputs a tool design that helps solve the task. A design-conditioned controller policy can then perform manipulation using these tools. In this work, we take a step towards this goal by introducing a reinforcement learning framework for jointly learning these policies. Through simulated manipulation tasks, we show that this framework is more sample efficient than prior methods in multi-goal or multi-variant settings, can perform zero-shot interpolation or fine-tuning to tackle previously unseen goals, and allows tradeoffs between the complexity of design and control policies under practical constraints. Finally, we deploy our learned policies onto a real robot. Please see our supplementary video and website at https://robotic-tool-design.github.io/ for visualizations.
翻译:受限于自身形态时,人类和某些动物物种展现出利用环境中的物体完成原本不可能实现的任务的非凡能力。机器人同样可通过工具使用解锁更多能力。当前通过深度学习联合优化形态与控制的先进技术,在设计运动智能体方面效果显著。然而,输出单一形态适用于运动任务,而操作任务需要根据当前目标采用多种策略。操作智能体必须能为不同目标快速生成专用工具。因此,我们提出学习一个设计策略(designer policy)而非单一设计。该设计策略以任务信息为条件,输出有助于解决该任务的工具设计。一个基于设计的控制器策略(design-conditioned controller policy)则可使用这些工具执行操作。本研究通过引入强化学习框架联合学习这些策略,向此目标迈进一步。通过模拟操作任务,我们证明该框架在多目标或多变体场景中样本效率优于先前方法,可进行零样本插值或微调以处理未见过的目标,并在实际约束下实现设计与控制策略复杂度之间的权衡。最后,我们将学习到的策略部署到真实机器人上。可视化结果请参见补充视频及网站 https://robotic-tool-design.github.io/。