Compared to general document analysis tasks, form document structure understanding and retrieval are challenging. Form documents are typically made by two types of authors; A form designer, who develops the form structure and keys, and a form user, who fills out form values based on the provided keys. Hence, the form values may not be aligned with the form designer's intention (structure and keys) if a form user gets confused. In this paper, we introduce Form-NLU, the first novel dataset for form structure understanding and its key and value information extraction, interpreting the form designer's intent and the alignment of user-written value on it. It consists of 857 form images, 6k form keys and values, and 4k table keys and values. Our dataset also includes three form types: digital, printed, and handwritten, which cover diverse form appearances and layouts. We propose a robust positional and logical relation-based form key-value information extraction framework. Using this dataset, Form-NLU, we first examine strong object detection models for the form layout understanding, then evaluate the key information extraction task on the dataset, providing fine-grained results for different types of forms and keys. Furthermore, we examine it with the off-the-shelf pdf layout extraction tool and prove its feasibility in real-world cases.
翻译:摘要:与通用文档分析任务相比,表单文档的结构理解与检索更具挑战性。表单文档通常由两类作者共同完成:表单设计者负责构建表单结构与标签键,而表单填写者则根据给定的标签键填写表单值。因此,若填写者对表单设计者的意图(即结构与标签键)产生混淆,表单值可能与设计者的预期不一致。本文提出Form-NLU——首个用于表单结构理解及其键值信息提取的新型数据集,该数据集旨在解读表单设计者的设计意图,并衡量用户填入值与之对齐的程度。该数据集包含857张表单图像、6000个表单键值对及4000个表格键值对,同时涵盖数字、印刷和手写三种表单类型,覆盖多样化的表单外观与布局。我们提出一种基于位置与逻辑关系的鲁棒表单键值信息提取框架。基于该Form-NLU数据集,我们首先评估了强目标检测模型在表单布局理解中的性能,随后在数据集上评价键信息提取任务,并针对不同表单类型与键类别给出细粒度结果。此外,我们采用现成的PDF布局提取工具对其进行验证,证明了该框架在实际场景中的可行性。