Optical Character Recognition is a technique that converts document images into searchable and editable text, making it a valuable tool for processing scanned documents. While the Farsi language stands as a prominent and official language in Asia, efforts to develop efficient methods for recognizing Farsi printed text have been relatively limited. This is primarily attributed to the languages distinctive features, such as cursive form, the resemblance between certain alphabet characters, and the presence of numerous diacritics and dot placement. On the other hand, given the substantial training sample requirements of deep-based architectures for effective performance, the development of such datasets holds paramount significance. In light of these concerns, this paper aims to present a novel large-scale dataset, IDPL-PFOD2, tailored for Farsi printed text recognition. The dataset comprises 2003541 images featuring a wide variety of fonts, styles, and sizes. This dataset is an extension of the previously introduced IDPL-PFOD dataset, offering a substantial increase in both volume and diversity. Furthermore, the datasets effectiveness is assessed through the utilization of both CRNN-based and Vision Transformer architectures. The CRNN-based model achieves a baseline accuracy rate of 78.49% and a normalized edit distance of 97.72%, while the Vision Transformer architecture attains an accuracy of 81.32% and a normalized edit distance of 98.74%.
翻译:光学字符识别是一种将文档图像转换为可搜索和可编辑文本的技术,使其成为处理扫描文档的重要工具。尽管波斯语作为亚洲一种重要且官方的语言,但针对波斯文印刷体文本识别的高效方法开发相对有限。这主要归因于该语言的独特特征,如连笔形式、某些字母字符的相似性,以及大量变音符号和点的存在。另一方面,鉴于基于深度学习的架构需要大量训练样本才能获得有效性能,此类数据集的开发至关重要。基于这些考虑,本文旨在提出一个面向波斯文印刷体文本识别的大规模新数据集IDPL-PFOD2。该数据集包含2003541张图像,涵盖多种字体、样式和字号。它是此前发布的IDPL-PFOD数据集的扩展,在数量和多样性上均有显著提升。此外,通过基于CRNN和视觉Transformer架构评估了数据集的有效性。基于CRNN的模型实现了78.49%的基线准确率和97.72%的归一化编辑距离,而视觉Transformer架构则达到了81.32%的准确率和98.74%的归一化编辑距离。