Multimodal reasoning is a critical component in the pursuit of artificial intelligence systems that exhibit human-like intelligence, especially when tackling complex tasks. While the chain-of-thought (CoT) technique has gained considerable attention, the existing ScienceQA dataset, which focuses on multimodal scientific questions and explanations from elementary and high school textbooks, lacks a comprehensive evaluation of diverse approaches. To address this gap, we present COCO Multi-Modal Reasoning Dataset(COCO-MMRD), a novel dataset that encompasses an extensive collection of open-ended questions, rationales, and answers derived from the large object dataset COCO. Unlike previous datasets that rely on multiple-choice questions, our dataset pioneers the use of open-ended questions in the context of multimodal CoT, introducing a more challenging problem that effectively assesses the reasoning capability of CoT models. Through comprehensive evaluations and detailed analyses, we provide valuable insights and propose innovative techniques, including multi-hop cross-modal attention and sentence-level contrastive learning, to enhance the image and text encoders. Extensive experiments demonstrate the efficacy of the proposed dataset and techniques, offering novel perspectives for advancing multimodal reasoning.
翻译:多模态推理是实现具备类人智能的人工智能系统的关键组成部分,尤其在处理复杂任务时更为重要。尽管思维链(CoT)技术已引起广泛关注,但现有聚焦于中小学教材多模态科学问题及解释的ScienceQA数据集,缺乏对不同方法的全面评估。为弥补这一空白,我们提出COCO多模态推理数据集(COCO-MMRD),这是一个涵盖来自大型目标检测数据集COCO的开放式问题、推理过程及答案的全新数据集。与以往依赖选择题的数据集不同,本数据集首次在多模态推理背景下采用开放式问题,引入更具挑战性的任务,从而有效评估CoT模型的推理能力。通过全面的评估与详细分析,我们获得了重要洞见,并提出创新技术——包括多跳跨模态注意力机制与句子级对比学习——以增强图像和文本编码器。大量实验证明了所提出数据集与技术的有效性,为推进多模态推理提供了全新视角。