The laws of model size, data volume, computation and model performance have been extensively studied in the field of Natural Language Processing (NLP). However, the scaling laws in Optical Character Recognition (OCR) have not yet been investigated. To address this, we conducted comprehensive studies that involved examining the correlation between performance and the scale of models, data volume and computation in the field of text recognition.Conclusively, the study demonstrates smooth power laws between performance and model size, as well as training data volume, when other influencing factors are held constant. Additionally, we have constructed a large-scale dataset called REBU-Syn, which comprises 6 million real samples and 18 million synthetic samples. Based on our scaling law and new dataset, we have successfully trained a scene text recognition model, achieving a new state-ofthe-art on 6 common test benchmarks with a top-1 average accuracy of 97.42%. The models and dataset are publicly available at https://github.com/large-ocr-model/large-ocr-model.github.io.
翻译:摘要:模型规模、数据量、计算量与模型性能之间的规律已在自然语言处理领域得到广泛研究,然而光学字符识别中的缩放定律尚未被探讨。为此,我们开展了系统性研究,旨在揭示文本识别领域中模型性能与规模、数据量及计算量之间的相关性。研究表明,在其他影响因素保持恒定时,模型性能与模型规模及训练数据量之间呈现平滑幂律关系。此外,我们构建了大规模数据集REBU-Syn,包含600万真实样本和1800万合成样本。基于提出的缩放定律与新数据集,我们成功训练了场景文本识别模型,在6个通用测试基准上取得了新的最优结果,Top-1平均准确率达97.42%。模型及数据集已开源发布于https://github.com/large-ocr-model/large-ocr-model.github.io。