Effective communication requires adapting to the idiosyncrasies of each communicative context--such as the common ground shared with each partner. Humans demonstrate this ability to specialize to their audience in many contexts, such as the popular game Dixit. We take inspiration from Dixit to formulate a multi-agent image reference game where a (trained) speaker model is rewarded for describing a target image such that one (pretrained) listener model can correctly identify it among distractors, but another listener cannot. To adapt, the speaker must exploit differences in the knowledge it shares with the different listeners. We show that finetuning an attention-based adapter between a CLIP vision encoder and a large language model in this contrastive, multi-agent setting gives rise to context-dependent natural language specialization from rewards only, without direct supervision. Through controlled experiments, we show that training a speaker with two listeners that perceive differently, using our method, allows the speaker to adapt to the idiosyncracies of the listeners. Furthermore, we show zero-shot transfer of the specialization to real-world data. Our experiments demonstrate a method for specializing grounded language models without direct supervision and highlight the interesting research challenges posed by complex multi-agent communication.
翻译:有效沟通需要适应每个沟通语境的特异性,例如与每个伙伴共享的共同基础。人类在许多语境中展示出这种专门化于受众的能力,例如流行的Dixit游戏。我们受Dixit启发,构建了一个多智能体图像指涉游戏,其中(经过训练的)说话者模型因描述目标图像而获得奖励,该描述能使一个(预训练的)听者模型正确识别目标图像(在干扰项中),而另一个听者则不能。为了适应,说话者必须利用其与不同听者共享知识的差异。我们表明,在CLIP视觉编码器与大语言模型之间微调基于注意力的适配器,在这种对比性、多智能体设置中,仅从奖励中即可产生依赖于语境的自然语言专门化,而无需直接监督。通过受控实验,我们证明,使用我们的方法训练一个具有两个感知不同的听者的说话者,能使说话者适应听者的特异性。此外,我们展示了该专门化到真实世界数据上的零样本迁移。我们的实验展示了一种无需直接监督即可专门化具有基础的语言模型的方法,并强调了由复杂多智能体通信带来的有趣研究挑战。