Parliamentary proceedings represent a rich yet challenging resource for computational analysis, particularly when preserved only as scanned historical documents. Existing efforts to transcribe Italian parliamentary speeches have relied on traditional Optical Character Recognition pipelines, resulting in transcription errors and limited semantic annotation. In this paper, we propose a pipeline based on Vision-Language Models for the automatic transcription, semantic segmentation, and entity linking of Italian parliamentary speeches. The pipeline employs a specialised OCR model to extract text while preserving reading order, followed by a large-scale Vision-Language Model that performs transcription refinement, element classification, and speaker identification by jointly reasoning over visual layout and textual content. Extracted speakers are then linked to the Chamber of Deputies knowledge base through SPARQL queries and a multi-strategy fuzzy matching procedure. Evaluation against an established benchmark demonstrates substantial improvements both in transcription quality and speaker tagging.
翻译:议会程序文件是计算分析的丰富资源,但若仅以扫描历史文档形式保存,则极具挑战性。现有意大利议会演讲转录工作依赖传统光学字符识别流水线,导致转录错误且语义标注有限。本文提出一种基于视觉语言模型的流水线,用于自动转录、语义分割和实体链接意大利议会演讲。该流水线首先采用专用OCR模型提取文本并保留阅读顺序,继而通过大规模视觉语言模型联合推理视觉版面和文本内容,完成转录精炼、元素分类和发言者识别。随后,通过SPARQL查询与多策略模糊匹配流程,将识别出的发言者链接至众议院知识库。在现有基准测试上的评估表明,该方法在转录质量和发言者标注方面均取得显著提升。