Natural language instruction following is paramount to enable collaboration between artificial agents and human beings. Natural language-conditioned reinforcement learning (RL) agents have shown how natural languages' properties, such as compositionality, can provide a strong inductive bias to learn complex policies. Previous architectures like HIGhER combine the benefit of language-conditioning with Hindsight Experience Replay (HER) to deal with sparse rewards environments. Yet, like HER, HIGhER relies on an oracle predicate function to provide a feedback signal highlighting which linguistic description is valid for which state. This reliance on an oracle limits its application. Additionally, HIGhER only leverages the linguistic information contained in successful RL trajectories, thus hurting its final performance and data-efficiency. Without early successful trajectories, HIGhER is no better than DQN upon which it is built. In this paper, we propose the Emergent Textual Hindsight Experience Replay (ETHER) agent, which builds on HIGhER and addresses both of its limitations by means of (i) a discriminative visual referential game, commonly studied in the subfield of Emergent Communication (EC), used here as an unsupervised auxiliary task and (ii) a semantic grounding scheme to align the emergent language with the natural language of the instruction-following benchmark. We show that the referential game's agents make an artificial language emerge that is aligned with the natural-like language used to describe goals in the BabyAI benchmark and that it is expressive enough so as to also describe unsuccessful RL trajectories and thus provide feedback to the RL agent to leverage the linguistic, structured information contained in all trajectories. Our work shows that EC is a viable unsupervised auxiliary task for RL and provides missing pieces to make HER more widely applicable.
翻译:自然语言指令跟随对于实现人工智能体与人类之间的协作至关重要。基于自然语言条件的强化学习智能体已展现出自然语言的组合性等特性如何为学习复杂策略提供强归纳偏置。HIGhER等先前架构将语言条件化与事后经验回放(HER)相结合,以应对稀疏奖励环境。然而,与HER类似,HIGhER依赖一个预言函数来提供反馈信号,指示哪些语言描述适用于哪些状态。这种对预言器的依赖限制了其应用范围。此外,HIGhER仅利用成功强化学习轨迹中的语言信息,从而损害了其最终性能和数据效率。若没有早期成功轨迹,HIGhER的性能甚至不如其基础模型DQN。本文提出新兴文本事后经验回放(ETHER)智能体,该智能体基于HIGhER构建,并通过以下方式解决上述两个局限:(i)利用新兴通信子领域常见的判别性视觉指代游戏作为无监督辅助任务;(ii)采用语义对齐机制将新兴语言与指令跟随基准中的自然语言对齐。我们证明,指代游戏中的智能体能够产生一种与BabyAI基准中描述目标所用的类自然语言对齐的人工语言,且该语言具备足够的表现力,可描述失败的强化学习轨迹,从而为强化学习智能体提供反馈,使其利用所有轨迹中包含的语言结构化信息。本研究表明,新兴通信是强化学习可行的无监督辅助任务,并为扩大HER的适用性提供了缺失环节。