Many templatized documents are programmatically generated from structured data following a visual template. Such documents include invoices, tax documents, financial reports, and purchase orders. Effective data extraction from these documents is crucial to support downstream analytical tasks. Current data extraction tools often struggle with complex document layouts, incur high latency and/or cost on large datasets, and require significant human effort. The key insight of our tool, TWIX, is to infer the underlying template used to create such documents, and then extract the data, rather than extracting directly from documents. To do so, TWIX first infers the underlying fields, such as columns of tabular portions or keys in co-located key-value pairs, by leveraging their consistent location patterns (e.g., two fields in the same template repeatedly co-occur within a fixed distance apart across multiple records). TWIX then assembles these fields into a template by enforcing visual constraints, such as vertically aligning table rows with their column headers for tabular regions, and horizontally aligning keys with their values for key-value pairs. TWIX then uses this inferred template to accurately and efficiently extract data from templatized documents at a low cost. On one benchmark with 34 diverse real-world datasets, TWIX outperforms state-of-the-art structured data extraction tools (Evaporate, Textract, and Azure Document Intelligence), and vision-based LLMs like GPT-4-Vision, by over 25% in precision and recall. Another benchmark with 30 large datasets demonstrates TWIX's scalability: it is 520X faster and 3,786X cheaper than the most competitive compared tool, for extracting data from large document collections with over 2000 pages.


翻译:许多模板化文档是根据结构化数据通过视觉模板程序化生成的。这类文档包括发票、税务文件、财务报表和采购订单。从这些文档中高效抽取数据对于支撑下游分析任务至关重要。当前的数据抽取工具在处理复杂文档布局时往往存在困难,在大规模数据集上会产生高延迟和/或高成本,且需要大量人工干预。我们的工具TWIX的核心思想在于:通过推理生成文档的底层模板来抽取数据,而非直接从文档中进行抽取。具体而言,TWIX首先利用字段间存在的一致位置模式(例如同一模板中的两个字段在多个记录中持续以固定距离共现)来推断底层字段,如表格部分的列或并列键值对中的键。随后,TWIX通过施加视觉约束将这些字段组装成模板——对表格区域垂直对齐表格行与列标题,对键值对水平对齐键与值。TWIX利用该推断出的模板,能够以低成本、高精度高效地从模板化文档中抽取数据。在包含34个多样化真实世界数据集的基准测试中,TWIX的精确率和召回率均比现有最先进的结构化数据抽取工具(Evaporate、Textract和Azure Document Intelligence)以及基于视觉的大语言模型(如GPT-4-Vision)高出25%以上。另一项包含30个大型数据集的基准测试证明了TWIX的可扩展性:在从超过2000页的大规模文档集合中抽取数据时,其速度是最具竞争力对比工具的520倍,成本仅为其1/3786。

0
下载
关闭预览

相关内容

文档视觉问答简述
专知会员服务
7+阅读 · 2025年10月17日
【博士论文】用于化学结构抽取的多模态文档理解
专知会员服务
9+阅读 · 2025年10月12日
【MIT博士论文】合成数据的视觉表示学习
专知会员服务
27+阅读 · 2024年8月25日
《基于深度学习的视觉文档信息抽取》研究综述
专知会员服务
36+阅读 · 2024年2月3日
专知会员服务
204+阅读 · 2020年10月14日
【关系抽取】从文本中进行关系抽取的几种不同的方法
深度学习自然语言处理
29+阅读 · 2020年3月30日
用深度学习做文本摘要
专知
24+阅读 · 2019年3月30日
计算机视觉方向简介 | 用深度学习进行表格提取
计算机视觉life
21+阅读 · 2019年2月19日
论文报告 | Graph-based Neural Multi-Document Summarization
科技创新与创业
15+阅读 · 2017年12月15日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
8+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Arxiv
0+阅读 · 6月15日
VIP会员
最新内容
从采集到决策:美军视角下的战术情报范式重构
专知会员服务
0+阅读 · 4分钟前
《履带式无人地面战车技术发展现状》
专知会员服务
0+阅读 · 刚刚
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
相关基金
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
8+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员