Multimodal information extraction (MIE) aims to extract structured information from unstructured multimedia content. Due to the diversity of tasks and settings, most current MIE models are task-specific and data-intensive, which limits their generalization to real-world scenarios with diverse task requirements and limited labeled data. To address these issues, we propose a novel multimodal question answering (MQA) framework to unify three MIE tasks by reformulating them into a unified span extraction and multi-choice QA pipeline. Extensive experiments on six datasets show that: 1) Our MQA framework consistently and significantly improves the performances of various off-the-shelf large multimodal models (LMM) on MIE tasks, compared to vanilla prompting. 2) In the zero-shot setting, MQA outperforms previous state-of-the-art baselines by a large margin. In addition, the effectiveness of our framework can successfully transfer to the few-shot setting, enhancing LMMs on a scale of 10B parameters to be competitive or outperform much larger language models such as ChatGPT and GPT-4. Our MQA framework can serve as a general principle of utilizing LMMs to better solve MIE and potentially other downstream multimodal tasks.
翻译:多模态信息抽取(MIE)旨在从非结构化多媒体内容中提取结构化信息。由于任务与设置的多样性,现有MIE模型多为任务特定型且依赖大量数据,这限制了它们在真实场景中应对多样化任务需求与有限标注数据时的泛化能力。为解决这些问题,我们提出一种新型多模态问答(MQA)框架,通过将三种MIE任务重构为统一的跨度抽取与多选问答流水线,实现任务统一。在六个数据集上的大量实验表明:1)与原始提示方法相比,我们的MQA框架能持续且显著提升各类现成大型多模态模型(LMM)在MIE任务上的性能;2)在零样本设定下,MQA大幅超越此前最优基线模型。此外,该框架的有效性可成功迁移至少样本设定,使参数规模达100亿的LMM能够与ChatGPT、GPT-4等更大规模语言模型竞争甚至超越其性能。我们的MQA框架可作为利用LMM更好解决MIE及其他潜在下游多模态任务的通用准则。