Missing-person and child-safety investigations rely on heterogeneous case documents, including structured forms, bulletin-style posters, and narrative web profiles. Variations in layout, terminology, and data quality impede rapid triage, large-scale analysis, and search-planning workflows. This paper introduces the Guardian Parser Pack, an AI-driven parsing and normalization pipeline that transforms multi-source investigative documents into a unified, schema-compliant representation suitable for operational review and downstream spatial modeling. The proposed system integrates (i) multi-engine PDF text extraction with Optical Character Recognition (OCR) fallback, (ii) rule-based source identification with source-specific parsers, (iii) schema-first harmonization and validation, and (iv) an optional Large Language Model (LLM)-assisted extraction pathway incorporating validator-guided repair and shared geocoding services. We present the system architecture, key implementation decisions, and output design, and evaluate performance using both gold-aligned extraction metrics and corpus-level operational indicators. On a manually aligned subset of 75 cases, the LLM-assisted pathway achieved substantially higher extraction quality than the deterministic comparator (F1 = 0.8664 vs. 0.2578), while across 517 parsed records per pathway it also improved aggregate key-field completeness (96.97\% vs. 93.23\%). The deterministic pathway remained much faster (mean runtime 0.03 s/record vs. 3.95 s/record for the LLM pathway). In the evaluated run, all LLM outputs passed initial schema validation, so validator-guided repair functioned as a built-in safeguard rather than a contributor to the observed gains. These results support controlled use of probabilistic AI within a schema-first, auditable pipeline for high-stakes investigative settings.


翻译:失踪人员和儿童安全调查依赖于异构的案件文档,包括结构化表格、公告式海报以及叙事性网络档案。布局、术语和数据质量的差异阻碍了快速分类、大规模分析和搜索规划工作流程。本文介绍了Guardian Parser Pack,这是一种由人工智能驱动的解析和标准化流水线,可将多源调查文档转化为符合统一模式规范的表示形式,适用于业务审查和下游空间建模。该系统集成了(i)多引擎PDF文本提取与光学字符识别(OCR)备用方案、(ii)基于规则的源识别与源专用解析器、(iii)模式优先的协调与验证,以及(iv)可选的基于大语言模型(LLM)的辅助抽取路径,该路径结合了验证器引导的修复与共享地理编码服务。我们展示了系统架构、关键实现决策和输出设计,并使用黄金对齐抽取指标和语料级业务指标评估了性能。在手动对齐的75个案例子集中,基于LLM的辅助路径相比确定性比较器实现了显著更高的抽取质量(F1=0.8664 vs. 0.2578),而在每条路径处理517条解析记录的规模上,它还将聚合关键字段完整度从93.23%提升至96.97%。确定性路径的速度明显更快(平均运行时间0.03秒/条记录,而LLM路径为3.95秒/条记录)。在评估运行中,所有LLM输出均通过了初始模式验证,因此验证器引导的修复功能充当了内置安全保障,而非观测到性能提升的直接贡献者。这些结果支持在高风险调查场景中,在可审计的、模式优先的流水线内对概率型AI进行受控使用。

0
下载
关闭预览

相关内容

大语言模型平台在国防情报应用中的对比
专知会员服务
18+阅读 · 4月22日
面向表格数据的大模型推理综述
专知会员服务
68+阅读 · 2023年12月26日
异常检测论文大列表:方法、应用、综述
专知
126+阅读 · 2019年7月15日
使用 Canal 实现数据异构
性能与架构
20+阅读 · 2019年3月4日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
44+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
VIP会员
最新内容
边缘计算的军事应用
专知会员服务
5+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
8+阅读 · 8月8日
《多域冲突比较支持模型》60页
专知会员服务
13+阅读 · 8月7日
相关VIP内容
大语言模型平台在国防情报应用中的对比
专知会员服务
18+阅读 · 4月22日
面向表格数据的大模型推理综述
专知会员服务
68+阅读 · 2023年12月26日
相关资讯
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
44+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员