Document-based Visual Question Answering examines the document understanding of document images in conditions of natural language questions. We proposed a new document-based VQA dataset, PDF-VQA, to comprehensively examine the document understanding from various aspects, including document element recognition, document layout structural understanding as well as contextual understanding and key information extraction. Our PDF-VQA dataset extends the current scale of document understanding that limits on the single document page to the new scale that asks questions over the full document of multiple pages. We also propose a new graph-based VQA model that explicitly integrates the spatial and hierarchically structural relationships between different document elements to boost the document structural understanding. The performances are compared with several baselines over different question types and tasks\footnote{The full dataset will be released after paper acceptance.
翻译:基于文档的视觉问答旨在通过自然语言问题考察文档图像的理解能力。我们提出了一个新的基于文档的VQA数据集——PDF-VQA,旨在从文档元素识别、文档布局结构理解、上下文理解以及关键信息提取等多个维度全面评估文档理解能力。该数据集突破了当前仅限于单页文档的理解规模,将问题范围扩展至多页完整文档。同时,我们提出了一种基于图的VQA模型,该模型显式整合了不同文档元素间的空间与层级结构关系,以增强文档结构理解能力。通过与多个基线模型在不同问题类型和任务上的性能对比,验证了该方法的有效性\footnote{完整数据集将在论文被接收后发布}。