In this work, we explore massive pre-training on synthetic word images for enhancing the performance on four benchmark downstream handwriting analysis tasks. To this end, we build a large synthetic dataset of word images rendered in several handwriting fonts, which offers a complete supervision signal. We use it to train a simple convolutional neural network (ConvNet) with a fully supervised objective. The vector representations of the images obtained from the pre-trained ConvNet can then be considered as encodings of the handwriting style. We exploit such representations for Writer Retrieval, Writer Identification, Writer Verification, and Writer Classification and demonstrate that our pre-training strategy allows extracting rich representations of the writers' style that enable the aforementioned tasks with competitive results with respect to task-specific State-of-the-Art approaches.
翻译:本研究探索在大规模合成单词图像上进行预训练,以提升四项基准下游手写分析任务的性能。为此,我们构建了一个使用多种手写字体渲染的大规模合成单词图像数据集,该数据集提供完整的监督信号。我们利用该数据集训练一个简单的卷积神经网络(ConvNet),其训练目标为完全监督学习。从预训练ConvNet中获得的图像向量表示可被视为手写风格的编码。我们将此类表示应用于书写者检索、书写者识别、书写者验证和书写者分类任务,并证明我们的预训练策略能够提取书写者风格的丰富表示,从而使得上述任务在性能上能与特定任务的最先进(State-of-the-Art)方法相竞争。