Document-to-LLM applications typically read uploaded PDFs by first translating them into text through a hidden extraction layer that users cannot observe or audit. We show that this layer enables split-view PDFs: one document can have two semantic views before model reasoning. By mining specification-permitted or implementation-tolerated representation gaps at the PDF render/extract boundary, we instantiate 25 extraction gaps (EG) in which extractors return attacker-controlled or extractor-dependent text while the rendered page shows benign or different content. The gaps form four families: semantic overrides, hidden semantic injection, reading-order splits, and font-decoding splits, and 14 gaps have no exact path/mechanism-level match in prior PDF-to-LLM attacks. We evaluate these gaps on 16 PDF processing stacks and 7 commercial LLM services. Each gap causes render-extract divergence on at least one stack. Under a gap-level exposure criterion, every evaluated service exposes at least one gap, with 12/25 to 21/25 exposed gaps. Exposure is driven mainly by the ingestion stack -- not model identity alone. We further show that tested safety filters cover only selected hidden-text constructions. To support triage, we develop a static screening scanner whose rules trigger on all 25 benchmark gaps, and discuss dual-view consistency as a longer-term defense direction.


翻译:文档到LLM的应用程序通常通过一个用户无法观察或审计的隐藏提取层,先将上传的PDF文件转换为文本,然后再进行读取。我们表明,这一层使得分视图PDF成为可能:一份文档在模型推理之前可以拥有两种语义视图。通过挖掘PDF渲染/提取边界上规范允许或实现容忍的表征间隙,我们具体实现了25种提取间隙(EG),其中提取器返回的是受攻击者控制或依赖于提取器的文本,而渲染页面显示的是良性的或不同的内容。这些间隙分为四类:语义覆盖、隐藏语义注入、阅读顺序分裂和字体解码分裂,其中14种间隙在先前的PDF到LLM攻击中不存在完全匹配的路径/机制层级。我们在16个PDF处理栈和7个商业LLM服务上评估了这些间隙。每种间隙至少在一个栈上引起了渲染与提取的差异。在间隙级暴露标准下,每个被评估的服务至少暴露了一个间隙,暴露间隙数量从12/25到21/25不等。暴露主要由摄取栈驱动——而不仅仅是模型身份。我们进一步表明,测试的安全过滤器仅覆盖了选定的隐藏文本构造。为了支持分类排查,我们开发了一个静态扫描筛查器,其规则可触发所有25个基准间隙,并讨论了双视图一致性作为长期防御方向。

0
下载
关闭预览

相关内容

利用多个大型语言模型:关于LLM集成的调研
专知会员服务
35+阅读 · 2025年2月27日
【AAAI2025】SAIL:面向样本的上下文学习用于文档信息提取
专知会员服务
21+阅读 · 2024年12月24日
《将大型语言模型(LLM)整合到海军作战规划中》
专知会员服务
132+阅读 · 2024年6月13日
【ICLR2024】能检测到LLM产生的错误信息吗?
专知会员服务
25+阅读 · 2024年1月23日
如何检测LLM内容?UCSB等最新首篇《LLM生成内容检测》综述
面试题:Word2Vec中为什么使用负采样?
七月在线实验室
46+阅读 · 2019年5月16日
面试题:文本摘要中的NLP技术
七月在线实验室
15+阅读 · 2019年5月13日
论文浅尝 | Global Relation Embedding for Relation Extraction
开放知识图谱
12+阅读 · 2019年3月3日
超像素、语义分割、实例分割、全景分割 傻傻分不清?
计算机视觉life
19+阅读 · 2018年11月27日
动态可视化指南:一步步拆解LSTM和GRU
论智
17+阅读 · 2018年10月25日
disentangled-representation-papers
CreateAMind
26+阅读 · 2018年9月12日
放弃 RNN/LSTM 吧,因为真的不好用!望周知~
人工智能头条
19+阅读 · 2018年4月24日
论文报告 | Graph-based Neural Multi-Document Summarization
科技创新与创业
15+阅读 · 2017年12月15日
基于 word2vec 和 CNN 的文本分类 :综述 & 实践
国家自然科学基金
1+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
18+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
印度精确打击与指挥架构的断层
专知会员服务
4+阅读 · 7月20日
美空军AI完成F-16战斗机自主空战历史性试飞
专知会员服务
5+阅读 · 7月20日
深入Project Maven:为何人工智能在战场上依然失灵
锻造未来士兵:外骨骼、基因工程与赛博格
专知会员服务
7+阅读 · 7月19日
《无人机蜂群通信技术研究》50页
专知会员服务
8+阅读 · 7月19日
相关资讯
面试题:Word2Vec中为什么使用负采样?
七月在线实验室
46+阅读 · 2019年5月16日
面试题:文本摘要中的NLP技术
七月在线实验室
15+阅读 · 2019年5月13日
论文浅尝 | Global Relation Embedding for Relation Extraction
开放知识图谱
12+阅读 · 2019年3月3日
超像素、语义分割、实例分割、全景分割 傻傻分不清?
计算机视觉life
19+阅读 · 2018年11月27日
动态可视化指南:一步步拆解LSTM和GRU
论智
17+阅读 · 2018年10月25日
disentangled-representation-papers
CreateAMind
26+阅读 · 2018年9月12日
放弃 RNN/LSTM 吧,因为真的不好用!望周知~
人工智能头条
19+阅读 · 2018年4月24日
论文报告 | Graph-based Neural Multi-Document Summarization
科技创新与创业
15+阅读 · 2017年12月15日
基于 word2vec 和 CNN 的文本分类 :综述 & 实践
相关基金
国家自然科学基金
1+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
18+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员