Visual commonsense understanding requires Vision Language (VL) models to not only understand image and text but also cross-reference in-between to fully integrate and achieve comprehension of the visual scene described. Recently, various approaches have been developed and have achieved high performance on visual commonsense benchmarks. However, it is unclear whether the models really understand the visual scene and underlying commonsense knowledge due to limited evaluation data resources. To provide an in-depth analysis, we present a Multimodal Evaluation (ME) pipeline to automatically generate question-answer pairs to test models' understanding of the visual scene, text, and related knowledge. We then take a step further to show that training with the ME data boosts the model's performance in standard VCR evaluation. Lastly, our in-depth analysis and comparison reveal interesting findings: (1) semantically low-level information can assist the learning of high-level information but not the opposite; (2) visual information is generally under utilization compared with text.
翻译:视觉常识理解要求视觉语言(VL)模型不仅能理解图像和文本,还需在两者之间进行交叉参照,以充分整合并实现对所描述视觉场景的理解。近年来,多种方法被提出并在视觉常识基准测试中取得了高性能。然而,由于评估数据资源有限,尚不清楚模型是否真正理解视觉场景及背后的常识知识。为进行深入分析,我们提出一个多模态评估(ME)流水线,可自动生成问答对,用于测试模型对视觉场景、文本及相关知识的理解。进一步,我们证明使用ME数据进行训练可提升模型在标准VCR评估中的性能。最后,深入分析与比较揭示了有趣发现:(1)语义层面的低层信息可辅助高层信息的学习,反之则不成立;(2)与文本相比,视觉信息普遍未被充分利用。