Extracting meaningful entities belonging to predefined categories from Visually-rich Form-like Documents (VFDs) is a challenging task. Visual and layout features such as font, background, color, and bounding box location and size provide important cues for identifying entities of the same type. However, existing models commonly train a visual encoder with weak cross-modal supervision signals, resulting in a limited capacity to capture these non-textual features and suboptimal performance. In this paper, we propose a novel \textbf{V}isually-\textbf{A}symmetric co\textbf{N}sisten\textbf{C}y \textbf{L}earning (\textsc{Vancl}) approach that addresses the above limitation by enhancing the model's ability to capture fine-grained visual and layout features through the incorporation of color priors. Experimental results on benchmark datasets show that our approach substantially outperforms the strong LayoutLM series baseline, demonstrating the effectiveness of our approach. Additionally, we investigate the effects of different color schemes on our approach, providing insights for optimizing model performance. We believe our work will inspire future research on multimodal information extraction.
翻译:从视觉丰富的表单类文档(VFDs)中提取属于预定义类别的高语义实体是一项具有挑战性的任务。视觉和布局特征(如字体、背景、颜色以及边界框的位置和大小)为识别相同类型的实体提供了重要线索。然而,现有模型通常使用弱跨模态监督信号训练视觉编码器,导致其捕获这些非文本特征的能力有限,性能欠佳。本文提出一种新型的视觉非对称一致性学习(Vancl)方法,通过引入颜色先验信息增强模型捕获细粒度视觉和布局特征的能力,从而解决上述局限性。在基准数据集上的实验结果表明,我们的方法显著优于强大的LayoutLM系列基线模型,验证了该方法的有效性。此外,我们研究了不同配色方案对方法的影响,为优化模型性能提供了见解。我们相信这项工作将启发未来对多模态信息提取的研究。