Properly defining a reward signal to efficiently train a reinforcement learning (RL) agent is a challenging task. Designing balanced objective functions from which a desired behavior can emerge requires expert knowledge, especially for complex environments. Learning rewards from human feedback or using large language models (LLMs) to directly provide rewards are promising alternatives, allowing non-experts to specify goals for the agent. However, black-box reward models make it difficult to debug the reward. In this work, we propose Object-Centric Assessment with Language Models (OCALM) to derive inherently interpretable reward functions for RL agents from natural language task descriptions. OCALM uses the extensive world-knowledge of LLMs while leveraging the object-centric nature common to many environments to derive reward functions focused on relational concepts, providing RL agents with the ability to derive policies from task descriptions.
翻译:为强化学习(RL)智能体高效训练而正确定义奖励信号是一项具有挑战性的任务。设计能产生期望行为的平衡目标函数需要专业知识,尤其是在复杂环境中。从人类反馈中学习奖励,或使用大型语言模型(LLMs)直接提供奖励,是两种有前景的替代方案,允许非专家为智能体指定目标。然而,黑盒奖励模型使得调试奖励变得困难。在本工作中,我们提出了基于语言模型的对象中心评估(OCALM),用于从自然语言任务描述中推导出本质上可解释的强化学习智能体奖励函数。OCALM利用LLMs的广泛世界知识,同时结合许多环境共有的对象中心特性,推导出专注于关系概念的奖励函数,从而赋予RL智能体从任务描述中推导策略的能力。