Document-based Visual Question Answering examines the document understanding of document images in conditions of natural language questions. We proposed a new document-based VQA dataset, PDF-VQA, to comprehensively examine the document understanding from various aspects, including document element recognition, document layout structural understanding as well as contextual understanding and key information extraction. Our PDF-VQA dataset extends the current scale of document understanding that limits on the single document page to the new scale that asks questions over the full document of multiple pages. We also propose a new graph-based VQA model that explicitly integrates the spatial and hierarchically structural relationships between different document elements to boost the document structural understanding. The performances are compared with several baselines over different question types and tasks\footnote{The full dataset will be released after paper acceptance.
翻译:文档视觉问答旨在通过自然语言问题考察对文档图像的理解能力。我们提出了一个新的文档型VQA数据集——PDF-VQA,从文档元素识别、文档布局结构理解、上下文理解及关键信息提取等多个维度全面检验文档理解能力。该数据集将当前局限于单页文档理解的规模扩展至多页完整文档的问答任务。同时,我们提出了一种新的基于图的VQA模型,该模型显式融合不同文档元素之间的空间与层次结构关系,以增强文档结构理解能力。我们针对不同问题类型和任务,与多个基线方法进行了性能对比。\footnote{论文录用后将发布完整数据集。}