Efforts over the past three decades have produced web archives containing billions of webpage snapshots and petabytes of data. The End of Term Web Archive alone contains, among other file types, millions of PDFs produced by the federal government. While preservation with web archives has been successful, significant challenges for access and discoverability remain. For example, current affordances for browsing the End of Term PDFs are limited to downloading and browsing individual PDFs, as well as performing basic keyword search across them. In this paper, we introduce GovScape, a public search system that supports multimodal searches across 10,015,993 federal government PDFs from the 2020 End of Term crawl (70,958,487 total PDF pages) - to our knowledge, all renderable PDFs in the 2020 crawl that are 50 pages or under. GovScape supports four primary forms of search over these 10 million PDFs: in addition to providing (1) filter conditions over metadata facets including domain and crawl date and (2) exact text search against the PDF text, we provide (3) semantic text search and (4) visual search against the PDFs across individual pages, enabling users to structure queries such as "redacted documents" or "pie charts." We detail the constituent components of GovScape, including the search affordances, embedding pipeline, system architecture, and open source codebase. Significantly, the total estimated compute cost for GovScape's pre-processing pipeline for 10 million PDFs was approximately $1,500, equivalent to 47,000 PDF pages per dollar spent on compute, demonstrating the potential for immediate scalability. Accordingly, we outline steps that we have already begun pursuing toward multimodal search at the 100+ million PDF scale. GovScape can be found at https://www.govscape.net.


翻译:过去三十年的努力已产生存储数十亿网页快照和PB级数据的网络存档。仅"任期终结网络存档"(End of Term Web Archive)就包含联邦政府生成的数百万份PDF文件。尽管网络存档在保存方面取得了成功,但在可访问性和可发现性方面仍面临重大挑战。例如,当前浏览任期终结PDF的功能仅限于下载和浏览单个PDF,以及执行基础的关键词搜索。本文介绍GovScape——一个支持多模态搜索的公共系统,涵盖2020年任期终结爬取数据中10,015,993份联邦政府PDF(总计70,958,487页PDF页面)——据我们所知,这是该次爬取中所有可渲染且页数不超过50页的PDF。GovScape为这1000万份PDF提供四种主要搜索形式:除提供(1)基于域名和爬取日期等元数据维度的过滤条件,以及(2)针对PDF文本的精确文本搜索外,我们还提供(3)语义文本搜索和(4)跨页面对PDF进行视觉搜索,使用户能够构建如"已编辑文档"或"饼图"等查询。我们详细阐述了GovScape的组成部分,包括搜索功能、嵌入流水线、系统架构和开源代码库。值得关注的是,GovScape为处理1000万份PDF的预处理流水线总计算成本估计约为1500美元,相当于每美元计算成本可处理47,000页PDF页面,展现了即时扩展的潜力。据此,我们概述了已开始推进的面向1亿+PDF规模的多模态搜索实施路径。GovScape可访问https://www.govscape.net。

0
下载
关闭预览

相关内容

《搜索型数据库白皮书》正式发布, 45页pdf
专知会员服务
34+阅读 · 2024年7月19日
【干货书】大数据小摘要,272页pdf,剑桥大学出版社
专知会员服务
42+阅读 · 2021年7月6日
【干货】20大推荐系统公共数据集分享
机器学习与推荐算法
68+阅读 · 2020年3月13日
20个安全可靠的免费数据源,各领域数据任你挑
机器学习算法与Python学习
14+阅读 · 2019年5月9日
Image Captioning 36页最新综述, 161篇参考文献
专知
90+阅读 · 2018年10月23日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
9+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
博士论文 | 面向大模型推理的内存高效算法
专知会员服务
0+阅读 · 今天15:20
美空军新型反无人机部队初探
专知会员服务
4+阅读 · 今天5:45
《防空交战流程的概率建模研究》
专知会员服务
6+阅读 · 今天5:04
ICML 2026 教程 | 数值优化理论还重要吗?
专知会员服务
4+阅读 · 7月26日
ICM 2026 | 陶哲轩:人工智能时代的数学
专知会员服务
8+阅读 · 7月26日
《反无人机交战场景下的战斗归零研究》
专知会员服务
7+阅读 · 7月26日
博士论文 | 用代码结构感知方法推进代码大模型
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
9+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员