For generative AI agents to partner effectively with human users, the ability to accurately predict human intent is critical. But this ability to collaborate remains limited by a critical deficit: an inability to model common ground. We present a referential communication experiment with a factorial design involving director-matcher pairs (human-human, human-AI, AI-human, and AI-AI) that interact with multiple turns in repeated rounds to match pictures of objects not associated with any obvious lexicalized labels. We show that LVLMs cannot interactively generate and resolve referring expressions in a way that enables smooth communication, a crucial skill that underlies human language use. We release our corpus of 356 dialogues (89 pairs over 4 rounds each) along with the online pipeline for data collection and the tools for analyzing accuracy, efficiency, and lexical overlap.
翻译:为了使生成式AI代理能够有效与人类用户协作,准确预测人类意图的能力至关重要。但这种协作能力仍因一个关键缺陷而受限:无法建模共同基础。我们提出了一项指涉交流实验,采用因子设计,涉及导演-匹配者配对(人类-人类、人类-AI、AI-人类、AI-AI),这些配对在多轮次中反复互动,以匹配不具有明显词汇化标签的物体图片。我们证明,LVLM无法以促进流畅交流的方式交互式生成和解析指涉表达,而这一关键能力正是人类语言使用的基础。我们发布了包含356段对话(89对,每对4轮)的语料库,以及用于数据收集的在线流程和分析准确性、效率及词汇重叠度的工具。