Document-based Visual Question Answering examines the document understanding of document images in conditions of natural language questions. We proposed a new document-based VQA dataset, PDF-VQA, to comprehensively examine the document understanding from various aspects, including document element recognition, document layout structural understanding as well as contextual understanding and key information extraction. Our PDF-VQA dataset extends the current scale of document understanding that limits on the single document page to the new scale that asks questions over the full document of multiple pages. We also propose a new graph-based VQA model that explicitly integrates the spatial and hierarchically structural relationships between different document elements to boost the document structural understanding. The performances are compared with several baselines over different question types and tasks\footnote{The full dataset will be released after paper acceptance.
翻译:基于文档的视觉问答研究通过自然语言问题考察文档图像的理解能力。我们提出了一个新的文档级VQA数据集——PDF-VQA,旨在从文档元素识别、文档布局结构理解、上下文理解及关键信息提取等多个维度全面评估文档理解能力。该数据集将当前局限于单页文档的理解规模,扩展至涵盖多页完整文档的问答任务。同时,我们提出了一种新的基于图的VQA模型,该模型显式整合了不同文档元素之间的空间与层级结构关系,以增强文档结构理解能力。通过对比多个基线模型在不同问题类型与任务上的表现,验证了所提方法的有效性\footnote{完整数据集将在论文接收后公开发布。}