Scientific articles published prior to the "age of digitization" in the late 1990s contain figures which are "trapped" within their scanned pages. While progress to extract figures and their captions has been made, there is currently no robust method for this process. We present a YOLO-based method for use on scanned pages, after they have been processed with Optical Character Recognition (OCR), which uses both grayscale and OCR-features. We focus our efforts on translating the intersection-over-union (IOU) metric from the field of object detection to document layout analysis and quantify "high localization" levels as an IOU of 0.9. When applied to the astrophysics literature holdings of the NASA Astrophysics Data System (ADS), we find F1 scores of 90.9% (92.2%) for figures (figure captions) with the IOU cut-off of 0.9 which is a significant improvement over other state-of-the-art methods.
翻译:在20世纪90年代末的"数字化时代"之前发表的科学文章,其图表被"困"在扫描页面中。尽管在提取图表及其图注方面已取得进展,但目前尚无稳健的方法。我们提出一种基于YOLO的方法,用于经过光学字符识别(OCR)处理的扫描页面,该方法同时利用灰度特征和OCR特征。我们重点将目标检测领域的交并比(IOU)指标迁移至文档布局分析领域,并将"高度局部化"水平量化为IOU=0.9。当应用于NASA天体物理数据系统(ADS)的天体物理学文献库时,在IOU截断值为0.9的条件下,图表(图注)的F1得分达到90.9%(92.2%),这显著优于其他现有方法。