Purpose: In this paper, we establish a baseline for handwritten stenography recognition, using the novel LION dataset, and investigate the impact of including selected aspects of stenographic theory into the recognition process. We make the LION dataset publicly available with the aim of encouraging future research in handwritten stenography recognition. Methods: A state-of-the-art text recognition model is trained to establish a baseline. Stenographic domain knowledge is integrated by applying four different encoding methods that transform the target sequence into representations, which approximate selected aspects of the writing system. Results are further improved by integrating a pre-training scheme, based on synthetic data. Results: The baseline model achieves an average test character error rate (CER) of 29.81% and a word error rate (WER) of 55.14%. Test error rates are reduced significantly by combining stenography-specific target sequence encodings with pre-training and fine-tuning, yielding CERs in the range of 24.5% - 26% and WERs of 44.8% - 48.2%. Conclusion: The obtained results demonstrate the challenging nature of stenography recognition. Integrating stenography-specific knowledge, in conjunction with pre-training and fine-tuning on synthetic data, yields considerable improvements. Together with our precursor study on the subject, this is the first work to apply modern handwritten text recognition to stenography. The dataset and our code are publicly available via Zenodo.
翻译:目的:本文利用新型LION数据集建立手写速记识别的基准,并研究在识别过程中融入速记理论的特定方面所产生的影响。我们公开提供LION数据集,旨在促进手写速记识别领域的未来研究。方法:训练当前最先进的文本识别模型以建立基准。通过应用四种不同的编码方法将速记领域知识整合其中,这些编码方法将目标序列转换为近似书写系统特定方面的表示。进一步通过基于合成数据的预训练方案提升结果。结果:基准模型在测试集上实现了29.81%的平均字符错误率(CER)和55.14%的词错误率(WER)。通过结合速记特定的目标序列编码与预训练和微调,测试错误率显著降低,CER达到24.5%–26%,WER达到44.8%–48.2%。结论:所得结果证明了速记识别的挑战性。融入速记特定知识并结合合成数据的预训练与微调,带来了显著的性能提升。与我们在该主题上的前期研究一道,这是首项将现代手写文本识别应用于速记的工作。数据集及代码已通过Zenodo公开提供。