Document collections of various domains, e.g., legal, medical, or financial, often share some underlying collection-wide structure, which captures information that can aid both human users and structure-aware models. We propose to identify the typical structure of document within a collection, which requires to capture recurring topics across the collection, while abstracting over arbitrary header paraphrases, and ground each topic to respective document locations. These requirements pose several challenges: headers that mark recurring topics frequently differ in phrasing, certain section headers are unique to individual documents and do not reflect the typical structure, and the order of topics can vary between documents. Subsequently, we develop an unsupervised graph-based method which leverages both inter- and intra-document similarities, to extract the underlying collection-wide structure. Our evaluations on three diverse domains in both English and Hebrew indicate that our method extracts meaningful collection-wide structure, and we hope that future work will leverage our method for multi-document applications and structure-aware models.
翻译:不同领域(如法律、医疗或金融)的文档集合通常共享某种潜在的集合级结构,这种结构能捕捉有助于人类用户和结构感知模型的信息。我们旨在识别集合内文档的典型结构,这需要捕捉集合中重复出现的主题,同时抽象化各种标题表述,并将每个主题映射到对应的文档位置。这些要求带来了若干挑战:标记重复主题的标题在措辞上常常不同,某些章节标题是单个文档特有的而不反映典型结构,且主题顺序在不同文档间可能变化。为此,我们提出了一种基于图的无监督方法,利用文档间和文档内的相似性来提取潜在的集合级结构。我们在英语和希伯来语三个不同领域的评估表明,我们的方法能提取有意义的集合级结构,并希望未来研究能将我们的方法用于多文档应用和结构感知模型。