Many digitized corpora suffer from low resources because annotations may be scarce, page scans are noisy and of poor resolution, or layouts are structurally complex in ways that negatively affect the quality of automatic transcription. Developing robust classification models for low-resource languages is inhibited by the lack of large-scale annotated data and by the frequent semantic complexity of page layouts. To this end, we have curated a complex-layout dataset, manually classified into eight distinct layout types based on their separator regions. To overcome data scarcity, we propose a novel training strategy in the form of a CNN-based classifier that employs strong, domain-aware augmentations to improve generalization. We utilize narrow anisotropic Gaussian masking to suppress incidental textual details while preserving essential separations, compelling the model to learn global geometric arrangements. Additionally, we implement reflection-induced label transformations to enrich the training distribution while maintaining label consistency across asymmetric categories. The results demonstrate that layout-specific augmentations can substantially improve page-level layout classification under severe annotation scarcity.
翻译:许多数字化语料库因标注稀缺、页面扫描件噪声大且分辨率低,或版面结构复杂而影响自动转录质量,从而面临低资源困境。低资源语言的鲁棒分类模型开发受限于大规模标注数据的缺失以及页面版面频繁出现的语义复杂性。为此,我们整理了一个复杂版面数据集,根据分隔区域特征将其人工划分为八种不同的版面类型。为应对数据匮乏问题,我们提出了一种基于CNN分类器的新颖训练策略,该策略采用强领域感知增强方法以提升泛化能力。我们利用窄向异性高斯掩膜抑制偶发性文本细节,同时保留关键分隔结构,迫使模型学习全局几何排布。此外,我们引入反射诱导的标签变换方法,在保持非对称类别标签一致性的同时丰富训练数据分布。实验结果表明,在标注极度稀缺条件下,版面特定增强方法能显著提升页面级版面分类性能。