Digital archiving is becoming widespread owing to its effectiveness in protecting valuable books and providing knowledge to many people electronically. In this paper, we propose a novel approach to leverage digital archives for machine learning. If we can fully utilize such digitized data, machine learning has the potential to uncover unknown insights and ultimately acquire knowledge autonomously, just like humans read books. As a first step, we design a dataset construction pipeline comprising an optical character reader (OCR), an object detector, and a layout analyzer for the autonomous extraction of image-text pairs. In our experiments, we apply our pipeline on old photo books to construct an image-text pair dataset, showing its effectiveness in image-text retrieval and insight extraction.
翻译:数字化归档因其在保护珍贵图书以及以电子形式向大众传播知识方面的有效性而日益普及。本文提出了一种利用数字档案进行机器学习的新方法。若能充分利用这类数字化数据,机器学习有望揭示未知见解,并最终像人类阅读书籍一样自主获取知识。作为第一步,我们设计了一个数据集构建流程,该流程包含光学字符识别器(OCR)、目标检测器和布局分析器,用于自主提取图文配对数据。实验中,我们将该流程应用于老旧摄影集,构建了一个图文配对数据集,并验证了其在图文检索及见解提取方面的有效性。