Structured documents--tables paired with captions, figures with explanations, equations with the paragraphs that interpret them--are routinely fragmented when indexed for retrieval. Element-level indexing treats every parsed element as an independent chunk, scattering semantically cohesive units across separate retrieval candidates. This paper presents a parser-independent pipeline that constructs Evidence Units (EUs): semantically complete document chunks that group visual assets with their contextual text. We introduce four contributions: (1) ontology-grounded role normalization extending DoCO that maps heterogeneous parser outputs to a unified semantic schema; (2) a semantic global assignment algorithm that optimally assigns paragraphs to EUs via a full similarity matrix; (3) a graph-based decision layer in Neo4j that formalizes EU construction rules and validates completeness through two invariants; and (4) cross-parser validation showing EU spatial footprints converge across MinerU and Docling, with gains preserved under parser-induced bbox variance. Experiments on OmniDocBench v1.0 (1,340 pages; 1,551 QA pairs) show EU-based chunking improves retrieval LCS by +0.31 (0.50 to 0.81). Recall@1 increases from 0.15 to 0.51 (3.4x) and MinK decreases from 2.58 to 1.72. Cross-parser results confirm the gain (LCS +0.23 to +0.31) is preserved across parsers. Text queries show the most dramatic gain: Recall@1 rises from 0.08 to 0.47.
翻译:结构化文档(如表格与标题配对、图表与说明文段、公式与解释段落)在索引检索时通常被碎片化处理。基于元素的索引将每个解析单元视为独立片段,导致语义完整的组织单元分散于不同检索候选结果中。本文提出一种与解析器无关的流水线,用于构建证据单元(EU):将视觉资源与其上下文文本组合为语义完整的文档片段。我们贡献四项创新:(1)基于本体驱动的角色规范化扩展DoCO,将异构解析器输出映射至统一语义框架;(2)语义全局分配算法,通过完整相似度矩阵将段落最优分配至证据单元;(3)基于Neo4j的图决策层,通过两个不变式形式化证据单元构建规则并验证完整性;(4)跨解析器验证表明,在MinerU与Docling生成的证据单元空间覆盖度收敛,且解析器边界框方差波动不影响增益效果。在OmniDocBench v1.0(1340页,1551个问答对)上的实验显示,基于证据单元的分块将检索LCS提升+0.31(从0.50增至0.81),Recall@1从0.15升至0.51(提升3.4倍),MinK从2.58降至1.72。跨解析器结果证实增益稳定性(LCS保持+0.23至+0.31),其中文本查询表现最为显著:Recall@1从0.08跃升至0.47。