Text segmentation, the task of dividing a document into sections, is often a prerequisite for performing additional natural language processing tasks. Existing text segmentation methods have typically been developed and tested using clean, narrative-style text with segments containing distinct topics. Here we consider a challenging text segmentation task: dividing newspaper marriage announcement lists into units of one announcement each. In many cases the information is not structured into sentences, and adjacent segments are not topically distinct from each other. In addition, the text of the announcements, which is derived from images of historical newspapers via optical character recognition, contains many typographical errors. As a result, these announcements are not amenable to segmentation with existing techniques. We present a novel deep learning-based model for segmenting such text and show that it significantly outperforms an existing state-of-the-art method on our task.
翻译:文本分割(将文档划分为不同部分的任务)通常是执行额外自然语言处理任务的前提。现有文本分割方法通常基于内容主题明确的段落式叙事文本进行开发和测试。本文研究一项具有挑战性的文本分割任务:将报纸结婚公告列表分割为单个公告单元。在许多情况下,这些信息并非以句子结构呈现,相邻段落之间也不存在主题差异。此外,通过光学字符识别从历史报纸图像中获取的公告文本包含大量印刷错误。这使得现有技术难以对这些公告进行有效分割。我们提出一种基于深度学习的新模型来处理此类文本分割,实验证明该模型在此任务上的性能显著优于当前最优方法。