Dialogue participants may have varying levels of knowledge about the topic under discussion. In such cases, it is essential for speakers to adapt their utterances by taking their audience into account. Yet, it is an open question how such adaptation can be modelled in computational agents. In this paper, we model a visually grounded referential game between a knowledgeable speaker and a listener with more limited visual and linguistic experience. Inspired by psycholinguistic theories, we endow our speaker with the ability to adapt its referring expressions via a simulation module that monitors the effectiveness of planned utterances from the listener's perspective. We propose an adaptation mechanism building on plug-and-play approaches to controlled language generation, where utterance generation is steered on the fly by the simulator without finetuning the speaker's underlying language model. Our results and analyses show that our approach is effective: the speaker's utterances become closer to the listener's domain of expertise, which leads to higher communicative success.
翻译:对话参与者对所讨论主题可能具有不同层次的知识。在这种情况下,说话者必须考虑听众来调整自己的话语。然而,如何在计算代理中建模这种适应性仍是一个待解决的问题。本文模拟了一个知识渊博的说话者与一个视觉和语言经验有限的听众之间基于视觉的指代游戏。受心理语言学理论的启发,我们为说话者赋予了一种能力,即通过一个模拟模块来调整其指代表达,该模块从听众的角度监控计划话语的有效性。我们提出了一种基于即插即用方法的可控语言生成适应机制,在该机制中,话语生成由模拟器即时引导,而无需微调说话者的底层语言模型。我们的结果表明,该方法有效:说话者的话语更贴近听众的专长领域,从而提高了沟通成功率。