Visually-situated languages such as charts and plots are omnipresent in real-world documents. These graphical depictions are human-readable and are often analyzed in visually-rich documents to address a variety of questions that necessitate complex reasoning and common-sense responses. Despite the growing number of datasets that aim to answer questions over charts, most only address this task in isolation, without considering the broader context of document-level question answering. Moreover, such datasets lack adequate common-sense reasoning information in their questions. In this work, we introduce a novel task named document-level chart question answering (DCQA). The goal of this task is to conduct document-level question answering, extracting charts or plots in the document via document layout analysis (DLA) first and subsequently performing chart question answering (CQA). The newly developed benchmark dataset comprises 50,010 synthetic documents integrating charts in a wide range of styles (6 styles in contrast to 3 for PlotQA and ChartQA) and includes 699,051 questions that demand a high degree of reasoning ability and common-sense understanding. Besides, we present the development of a potent question-answer generation engine that employs table data, a rich color set, and basic question templates to produce a vast array of reasoning question-answer pairs automatically. Based on DCQA, we devise an OCR-free transformer for document-level chart-oriented understanding, capable of DLA and answering complex reasoning and common-sense questions over charts in an OCR-free manner. Our DCQA dataset is expected to foster research on understanding visualizations in documents, especially for scenarios that require complex reasoning for charts in the visually-rich document. We implement and evaluate a set of baselines, and our proposed method achieves comparable results.
翻译:视觉情境化语言(如图表与绘图)在真实文档中无处不在。这些图形化表示具有人类可读性,常被用于视觉丰富型文档的分析中,以回答需要复杂推理与常识性反馈的各类问题。尽管现有图表问答数据集数量不断增长,但多数仅孤立地处理该任务,未考虑文档级问答的宏观语境。此外,这类数据集的问句缺乏充分的常识推理信息。本文提出一项新任务——文档级图表问答(DCQA)。该任务目标为:首先通过文档布局分析(DLA)提取文档中的图表,随后执行图表问答(CQA),最终实现文档级问答。我们构建的全新基准数据集包含50,010份合成文档,整合了多种风格的图表(6种风格,对比PlotQA与ChartQA的3种),并包含699,051个需要高推理能力与常识理解的问句。此外,我们开发了强大的问答生成引擎,利用表格数据、丰富色彩集及基础问句模板,自动生成大量推理型问答对。基于DCQA,我们设计了面向文档级图表理解的免OCR Transformer模型,能够以免OCR方式完成文档布局分析及复杂推理与常识问答。DCQA数据集有望推动文档可视化理解的研究,尤其针对需复杂推理的视觉丰富型文档场景。我们实现并评估了多组基线方法,所提方法取得了可比结果。