Two recent studies (Jones et al. (2026); Zeng et al. (2026)) reach apparently contradictory conclusions about whether LVLMs can coordinate on efficient referring expressions. We control for task differences between the studies while directly comparing their prompting styles. We replicate the finding that models can coordinate efficient referring expressions when explicitly prompted to do so, suggesting that other task differences are not responsible for divergent results. However, we also find that the same models fail to infer the need for communicative efficiency from a more implicit prompt, highlighting critical differences between how humans and AI systems communicate.
翻译:两项近期研究(Jones 等, 2026;Zeng 等, 2026)关于LVLM能否协调出高效指称表达式得出了看似矛盾的结论。我们在控制研究间任务差异的同时,直接比较了它们的提示风格。我们复现了如下发现:当被显式提示时,模型能够协调出高效的指称表达式,这表明其他任务差异并非导致结果分歧的原因。然而,我们也发现,相同的模型未能从更隐式的提示中推断出沟通效率的需求,这凸显了人类与AI系统在沟通方式上的关键差异。