We introduce IRIS (Intent Resolution via Inference-time Saccades), a novel training-free approach that uses eye-tracking data in real-time to resolve ambiguity in open-ended VQA. Through a comprehensive user study with 500 unique image-question pairs, we demonstrate that fixations closest to the time participants start verbally asking their questions are the most informative for disambiguation in Large VLMs, more than doubling the accuracy of responses on ambiguous questions (from 35.2% to 77.2%) while maintaining performance on unambiguous queries. We evaluate our approach across state-of-the-art VLMs, showing consistent improvements when gaze data is incorporated in ambiguous image-question pairs, regardless of architectural differences. We release a new benchmark dataset to use eye movement data for disambiguated VQA, a novel real-time interactive protocol, and an evaluation suite.
翻译:我们提出IRIS(推理时扫视意图消解),一种基于实时眼动数据的新型免训练方法,用于消解开放式视觉问答中的歧义。通过包含500个独特图文对的综合性用户研究,我们发现:在大型视觉语言模型中,参与者开始口头提问时刻最近的注视点对消歧最具信息量,可将模糊问题的响应准确率提升一倍以上(从35.2%提高至77.2%),同时保持对非模糊查询的性能。我们跨越多个最先进视觉语言模型进行评估,结果表明在模糊图文对中融入注视数据后,无论架构差异如何,均获得一致改进。我们发布了用于歧义消解VQA的眼动数据集基准、新型实时交互协议及评估套件。