In many machine learning tasks, a large general dataset and a small specialized dataset are available. In such situations, various domain adaptation methods can be used to adapt a general model to the target dataset. We show that in the case of neural networks trained for handwriting recognition using CTC, simple finetuning with data augmentation works surprisingly well in such scenarios and that it is resistant to overfitting even for very small target domain datasets. We evaluated the behavior of finetuning with respect to augmentation, training data size, and quality of the pre-trained network, both in writer-dependent and writer-independent settings. On a large real-world dataset, finetuning provided an average relative CER improvement of 25 % with 16 text lines for new writers and 50 % for 256 text lines.
翻译:在许多机器学习任务中,通常可获得一个大型通用数据集和一个小型专业数据集。在此类场景下,可采用多种领域自适应方法将通用模型适配至目标数据集。我们研究发现,在使用CTC损失函数训练的手写识别神经网络中,结合数据增强的简单微调方法在类似场景下表现出惊人的有效性,即使在目标域数据集规模极小的情况下也能有效抵抗过拟合。我们在作者依赖与作者独立两种设定下,系统评估了微调方法关于数据增强、训练数据规模及预训练网络质量的行为特性。在真实大规模数据集上,针对新书写者,16行文本可实现平均相对字符错误率(CER)降低25%,而256行文本则可达到50%的相对改善率。