Document-based Visual Question Answering examines the document understanding of document images in conditions of natural language questions. We proposed a new document-based VQA dataset, PDF-VQA, to comprehensively examine the document understanding from various aspects, including document element recognition, document layout structural understanding as well as contextual understanding and key information extraction. Our PDF-VQA dataset extends the current scale of document understanding that limits on the single document page to the new scale that asks questions over the full document of multiple pages. We also propose a new graph-based VQA model that explicitly integrates the spatial and hierarchically structural relationships between different document elements to boost the document structural understanding. The performances are compared with several baselines over different question types and tasks\footnote{The full dataset will be released after paper acceptance.
翻译:基于文档的视觉问答通过自然语言问题考察文档图像的文档理解能力。我们提出了一个新的基于文档的VQA数据集——PDF-VQA,旨在从多个维度全面评估文档理解能力,包括文档元素识别、文档布局结构理解、上下文理解以及关键信息提取。该数据集将当前局限于单页的文档理解规模扩展至跨多页完整文档的问答任务。此外,我们提出了一种新型基于图的VQA模型,显式融合不同文档元素间的空间与层级结构关系,以增强文档结构理解。针对不同问题类型与任务,我们与多个基线模型进行了性能对比\footnote{完整数据集将在论文接收后公开发布。}