Document dewarping from a distorted camera-captured image is of great value for OCR and document understanding. The document boundary plays an important role which is more evident than the inner region in document dewarping. Current learning-based methods mainly focus on complete boundary cases, leading to poor document correction performance of documents with incomplete boundaries. In contrast to these methods, this paper proposes MataDoc, the first method focusing on arbitrary boundary document dewarping with margin and text aware regularizations. Specifically, we design the margin regularization by explicitly considering background consistency to enhance boundary perception. Moreover, we introduce word position consistency to keep text lines straight in rectified document images. To produce a comprehensive evaluation of MataDoc, we propose a novel benchmark ArbDoc, mainly consisting of document images with arbitrary boundaries in four typical scenarios. Extensive experiments confirm the superiority of MataDoc with consideration for the incomplete boundary on ArbDoc and also demonstrate the effectiveness of the proposed method on DocUNet, DIR300, and WarpDoc datasets.
翻译:从畸变的相机拍摄图像中进行文档去扭曲对OCR和文档理解具有重要价值。文档边界在去扭曲过程中比内部区域更为显著,对校正效果起关键作用。现有基于学习方法主要针对完整边界场景,导致对不完整边界文档的校正性能不佳。与之不同,本文提出MataDoc——首个专注于任意边界文档去扭曲的方法,引入边缘感知与文本感知正则化。具体而言,通过显式考虑背景一致性设计边缘正则化以增强边界感知能力;同时,引入词位置一致性约束,确保校正后文档图像中的文本行保持平直。为全面评估MataDoc,我们构建了新型基准数据集ArbDoc,主要包含四种典型场景下的任意边界文档图像。大量实验表明,MataDoc在处理ArbDoc数据集中的不完整边界时具有优越性,并在DocUNet、DIR300和WarpDoc数据集上验证了所提方法的有效性。