Restoring the original, flat appearance of a printed document from casual photographs of bent and wrinkled pages is a common everyday problem. In this paper we propose a novel method for grid-based single-image document unwarping. Our method performs geometric distortion correction via a deep fully convolutional neural network that learns to predict the 3D grid mesh of the document and the corresponding 2D unwarping grid in a multi-task fashion, implicitly encoding the coupling between the shape of a 3D object and its 2D image. We additionally create and publish our own dataset, called UVDoc, which combines pseudo-photorealistic document images with ground truth grid-based physical 3D and unwarping information, allowing unwarping models to train on data that is more realistic in appearance than the commonly used synthetic Doc3D dataset, whilst also being more physically accurate. Our dataset is labeled with all the information necessary to train our unwarping network, without having to engineer separate loss functions that can deal with the lack of ground-truth typically found in document in the wild datasets. We include a thorough evaluation that demonstrates that our dual-task unwarping network trained on a mix of synthetic and pseudo-photorealistic images achieves state-of-the-art performance on the DocUNet benchmark dataset. Our code, results and UVDoc dataset will be made publicly available upon publication.
翻译:从随意拍摄的弯曲褶皱文档照片中恢复其原始平整外观是日常生活中的常见问题。本文提出一种基于网格的单图像文档展开新方法。该方法通过深度全卷积神经网络实现几何畸变校正,该网络以多任务学习方式同时预测文档的3D网格模型与对应的2D展开网格,隐式编码了三维物体形状与其二维图像之间的耦合关系。我们额外创建并发布了名为UVDoc的自建数据集,该数据集将伪真实感文档图像与基于网格的物理三维及展开信息的真值相结合,使得展开模型能在比常用合成Doc3D数据集更具视觉真实感且物理精度更高的数据上进行训练。该数据集标注了训练展开网络所需的全部信息,避免了针对野外文档数据集缺乏真值问题而设计独立损失函数的困扰。通过全面评估表明,在合成与伪真实感图像混合数据上训练的双任务展开网络在DocUNet基准数据集上达到了最优性能。我们的代码、结果及UVDoc数据集将在发表后公开。