In the real world, knowledge often exists in a multimodal and heterogeneous form. Addressing the task of question answering with hybrid data types, including text, tables, and images, is a challenging task (MMHQA). Recently, with the rise of large language models (LLM), in-context learning (ICL) has become the most popular way to solve QA problems. We propose MMHQA-ICL framework for addressing this problems, which includes stronger heterogeneous data retriever and an image caption module. Most importantly, we propose a Type-specific In-context Learning Strategy for MMHQA, enabling LLMs to leverage their powerful performance in this task. We are the first to use end-to-end LLM prompting method for this task. Experimental results demonstrate that our framework outperforms all baselines and methods trained on the full dataset, achieving state-of-the-art results under the few-shot setting on the MultimodalQA dataset.
翻译:现实世界中,知识常以多模态异质形式存在。处理包含文本、表格与图像的混合数据类型问答任务(MMHQA)极具挑战性。近年来,随着大语言模型(LLM)的兴起,情境学习(ICL)已成为解决问答问题的主流方法。我们提出MMHQA-ICL框架以应对该问题,该框架包含更强的异质数据检索器与图像描述模块。尤为重要的是,我们提出面向MMHQA的类型特异性情境学习策略,使LLM能够在此任务中发挥其强大性能。我们是首个将该任务作为端到端LLM提示方法的工作。实验结果表明,我们的框架优于所有基线方法及在全数据集上训练的模型,在MultimodalQA数据集的小样本设置下取得了最先进的结果。