Document-to-LLM applications typically read uploaded PDFs by first translating them into text through a hidden extraction layer that users cannot observe or audit. We show that this layer enables split-view PDFs: one document can have two semantic views before model reasoning. By mining specification-permitted or implementation-tolerated representation gaps at the PDF render/extract boundary, we instantiate 25 extraction gaps (EG) in which extractors return attacker-controlled or extractor-dependent text while the rendered page shows benign or different content. The gaps form four families: semantic overrides, hidden semantic injection, reading-order splits, and font-decoding splits, and 14 gaps have no exact path/mechanism-level match in prior PDF-to-LLM attacks. We evaluate these gaps on 16 PDF processing stacks and 7 commercial LLM services. Each gap causes render-extract divergence on at least one stack. Under a gap-level exposure criterion, every evaluated service exposes at least one gap, with 12/25 to 21/25 exposed gaps. Exposure is driven mainly by the ingestion stack -- not model identity alone. We further show that tested safety filters cover only selected hidden-text constructions. To support triage, we develop a static screening scanner whose rules trigger on all 25 benchmark gaps, and discuss dual-view consistency as a longer-term defense direction.
翻译:文档到LLM的应用程序通常通过一个用户无法观察或审计的隐藏提取层,先将上传的PDF文件转换为文本,然后再进行读取。我们表明,这一层使得分视图PDF成为可能:一份文档在模型推理之前可以拥有两种语义视图。通过挖掘PDF渲染/提取边界上规范允许或实现容忍的表征间隙,我们具体实现了25种提取间隙(EG),其中提取器返回的是受攻击者控制或依赖于提取器的文本,而渲染页面显示的是良性的或不同的内容。这些间隙分为四类:语义覆盖、隐藏语义注入、阅读顺序分裂和字体解码分裂,其中14种间隙在先前的PDF到LLM攻击中不存在完全匹配的路径/机制层级。我们在16个PDF处理栈和7个商业LLM服务上评估了这些间隙。每种间隙至少在一个栈上引起了渲染与提取的差异。在间隙级暴露标准下,每个被评估的服务至少暴露了一个间隙,暴露间隙数量从12/25到21/25不等。暴露主要由摄取栈驱动——而不仅仅是模型身份。我们进一步表明,测试的安全过滤器仅覆盖了选定的隐藏文本构造。为了支持分类排查,我们开发了一个静态扫描筛查器,其规则可触发所有25个基准间隙,并讨论了双视图一致性作为长期防御方向。