Generating robot demonstrations through simulation is widely recognized as an effective way to scale up robot data. Previous work often trained reinforcement learning agents to generate expert policies, but this approach lacks sample efficiency. Recently, a line of work has attempted to generate robot demonstrations via differentiable simulation, which is promising but heavily relies on reward design, a labor-intensive process. In this paper, we propose DiffGen, a novel framework that integrates differentiable physics simulation, differentiable rendering, and a vision-language model to enable automatic and efficient generation of robot demonstrations. Given a simulated robot manipulation scenario and a natural language instruction, DiffGen can generate realistic robot demonstrations by minimizing the distance between the embedding of the language instruction and the embedding of the simulated observation after manipulation. The embeddings are obtained from the vision-language model, and the optimization is achieved by calculating and descending gradients through the differentiable simulation, differentiable rendering, and vision-language model components, thereby accomplishing the specified task. Experiments demonstrate that with DiffGen, we could efficiently and effectively generate robot data with minimal human effort or training time.
翻译:通过仿真生成机器人演示被广泛认为是扩展机器人数据的有效方式。以往工作通常训练强化学习智能体来生成专家策略,但这种方法样本效率较低。近期,有研究尝试通过可微仿真生成机器人演示,这一方法虽然前景可观,却严重依赖奖励设计这一劳动密集型过程。本文提出DiffGen这一新型框架,它整合了可微物理仿真、可微渲染与视觉语言模型,能够自动、高效地生成机器人演示。在给定仿真机器人操作场景和自然语言指令的条件下,DiffGen通过最小化语言指令嵌入与操作后仿真观测嵌入之间的距离来生成逼真的机器人演示。嵌入由视觉语言模型获得,优化则通过计算并沿可微仿真、可微渲染及视觉语言模型组件传播梯度来实现,从而完成指定任务。实验表明,使用DiffGen能以最少的人力和训练时间高效生成机器人数据。