We propose a novel task, G4C, to study teacher-student natural language interactions in a goal-driven and grounded environment. Dungeons and Dragons (D&D), a role-playing game, provides an ideal setting to investigate such interactions. Here, the Dungeon Master (DM), i.e., the teacher, guides the actions of several players -- students, each with their own personas and abilities -- to achieve shared goals grounded in a fantasy world. Our approach is to decompose and model these interactions into (1) the DM's intent to guide players toward a given goal; (2) the DM's guidance utterance to the players expressing this intent; and (3) a theory-of-mind (ToM) model that anticipates the players' reaction to the guidance one turn into the future. We develop a novel reinforcement learning (RL) method for training a DM that generates guidance for players by rewarding utterances where the intent matches the ToM-anticipated player actions. Human and automated evaluations show that a DM trained to explicitly model intents and incorporate ToM of the players using RL generates better-quality guidance that is 3x more likely to fulfill the DM's intent than a vanilla natural language generation (NLG) approach.
翻译:摘要:本文提出一项新颖任务G4C,旨在研究目标驱动且具身环境下的师生自然语言交互。角色扮演游戏《龙与地下城》(D&D)为探究此类交互提供了理想场景。在该游戏中,地下城主(DM)作为教师,引导多名拥有各自人设与能力的玩家(学生)采取行动,以实现共同的目标,这些目标根植于一个奇幻世界。我们的方法将这些交互分解并建模为:(1)DM引导玩家走向特定目标的意图;(2)DM向玩家表达该意图的引导话语;(3)一个心智理论(ToM)模型,用于预测玩家在下一轮对引导的反应。我们开发了一种新颖的强化学习(RL)方法训练DM,通过奖励那些意图与ToM预测的玩家行为相匹配的话语,来生成针对玩家的引导。人工与自动评估表明,使用RL训练的、显式建模意图并融入玩家ToM的DM,能够生成更高质量的引导,其实现DM意图的可能性比普通的自然语言生成(NLG)方法高出3倍。