Automated transcription of parliamentary proceedings faces significant hurdles due to demographic bias, dialectal variation, and technical artifacts such as utterance truncation during segmentation. This paper introduces the ROManian PARliamentary Speech Corpus (ROMPAR) dataset, a 17.80-hour corpus of Romanian and Moldavian parliamentary speech, featuring double-annotated ground truth and explicit labels for reconstructed word fragments. To build a robust ASR system, we propose a multi-task adversarial training framework that enforces demographic invariance across age, gender, and dialect. We address the inherent instability of adversarial objectives in generative architectures by introducing an exponential decay mechanism for the adversarial coefficients. Furthermore, we implement an LLM-guided decoding strategy with position-dependent weighting to facilitate morphological completion of truncated terminal words. Our results demonstrate that the proposed framework significantly reduces WER and achieves an F1-score of 96.6% in morphological reconstruction.
翻译:摘要:自动转录议会会议记录面临显著挑战,包括人口统计偏差、方言变异以及分割过程中话语截断等技术伪影。本文介绍了罗马尼亚议会语音语料库ROMPAR数据集,这是一个17.80小时的罗马尼亚语和摩尔达维亚语议会语音语料库,包含双重标注的真实标注数据和重建词片段的显式标签。为构建鲁棒的自动语音识别系统,我们提出了一种多任务对抗训练框架,在年龄、性别和方言维度上实现人口统计不变性。针对生成式架构中对抗目标固有的不稳定性,我们引入了一种对抗系数的指数衰减机制。此外,我们实现了一种结合位置依赖权重的LLM引导解码策略,以促进截断终端词的形态补全。实验结果表明,所提框架显著降低了词错误率,并在形态重建上实现了96.6%的F1分数。