Visual question answering (VQA) has the potential to make the Internet more accessible in an interactive way, allowing people who cannot see images to ask questions about them. However, multiple studies have shown that people who are blind or have low-vision prefer image explanations that incorporate the context in which an image appears, yet current VQA datasets focus on images in isolation. We argue that VQA models will not fully succeed at meeting people's needs unless they take context into account. To further motivate and analyze the distinction between different contexts, we introduce Context-VQA, a VQA dataset that pairs images with contexts, specifically types of websites (e.g., a shopping website). We find that the types of questions vary systematically across contexts. For example, images presented in a travel context garner 2 times more "Where?" questions, and images on social media and news garner 2.8 and 1.8 times more "Who?" questions than the average. We also find that context effects are especially important when participants can't see the image. These results demonstrate that context affects the types of questions asked and that VQA models should be context-sensitive to better meet people's needs, especially in accessibility settings.
翻译:视觉问答(VQA)有望通过交互方式提升互联网的可访问性,帮助无法看到图像的用户提出相关问题。然而,多项研究表明,盲人或低视力群体更倾向于结合图像所在情境的解释,而现有VQA数据集仅关注孤立图像。我们认为,若不纳入上下文情境,VQA模型将无法完全满足用户需求。为深入探讨并分析不同情境的差异性,我们提出了Context-VQA——一个将图像与特定网站类型(如购物网站)等情境配对的VQA数据集。研究发现,问题类型随情境系统性地变化:例如,旅行情境中的图像引发"在哪里?"问题的频率是平均值的2倍,社交媒体和新闻情境中"是谁?"问题的频率分别达平均值的2.8倍和1.8倍。此外,当参与者无法看到图像时,情境效应尤为显著。这些结果表明,情境会影响问题的类型,而VQA模型应具备情境感知能力以更好地满足用户需求,尤其在可访问性场景中。