Large Language Models (LLMs) have excelled as high-level semantic planners for sequential decision-making tasks. However, harnessing them to learn complex low-level manipulation tasks, such as dexterous pen spinning, remains an open problem. We bridge this fundamental gap and present Eureka, a human-level reward design algorithm powered by LLMs. Eureka exploits the remarkable zero-shot generation, code-writing, and in-context improvement capabilities of state-of-the-art LLMs, such as GPT-4, to perform evolutionary optimization over reward code. The resulting rewards can then be used to acquire complex skills via reinforcement learning. Without any task-specific prompting or pre-defined reward templates, Eureka generates reward functions that outperform expert human-engineered rewards. In a diverse suite of 29 open-source RL environments that include 10 distinct robot morphologies, Eureka outperforms human experts on 83% of the tasks, leading to an average normalized improvement of 52%. The generality of Eureka also enables a new gradient-free in-context learning approach to reinforcement learning from human feedback (RLHF), readily incorporating human inputs to improve the quality and the safety of the generated rewards without model updating. Finally, using Eureka rewards in a curriculum learning setting, we demonstrate for the first time, a simulated Shadow Hand capable of performing pen spinning tricks, adeptly manipulating a pen in circles at rapid speed.
翻译:大型语言模型(LLMs)在顺序决策任务中已展现出作为高层语义规划器的卓越能力。然而,如何利用它们学习复杂的低层操作任务,例如灵巧的笔旋转,仍是一个未解难题。我们弥合了这一根本差距,并提出Eureka——一种由LLMs驱动的人类级奖励设计算法。Eureka利用最先进LLMs(如GPT-4)显著的零样本生成、代码编写和上下文内改进能力,对奖励代码执行进化优化。由此产生的奖励可通过强化学习用于获取复杂技能。无需特定任务提示或预定义奖励模板,Eureka即可生成优于人类专家设计的奖励函数。在包含10种不同机器人形态的29个开源强化学习环境套件中,Eureka在83%的任务上超越人类专家,平均标准化提升达52%。Eureka的通用性还催生了一种基于无梯度上下文内学习的强化学习从人类反馈方法,无需模型更新即可直接纳入人类输入以改善生成奖励的质量与安全性。最后,通过在课程学习场景中使用Eureka奖励,我们首次展示了模拟Shadow Hand执行笔旋转技巧——以快速旋转动作娴熟操控笔杆在圆轨迹上运动。