Recent advancements in Named Entity Recognition (NER) have significantly improved the identification of entities in textual data. However, spoken NER, a specialized field of spoken document retrieval, lags behind due to its limited research and scarce datasets. Moreover, cross-lingual transfer learning in spoken NER has remained unexplored. This paper utilizes transfer learning across Dutch, English, and German using pipeline and End-to-End (E2E) schemes. We employ Wav2Vec2-XLS-R models on custom pseudo-annotated datasets and investigate several architectures for the adaptability of cross-lingual systems. Our results demonstrate that End-to-End spoken NER outperforms pipeline-based alternatives over our limited annotations. Notably, transfer learning from German to Dutch surpasses the Dutch E2E system by 7% and the Dutch pipeline system by 4%. This study not only underscores the feasibility of transfer learning in spoken NER but also sets promising outcomes for future evaluations, hinting at the need for comprehensive data collection to augment the results.
翻译:近期命名实体识别技术的进步显著提升了文本数据中的实体识别能力。然而,作为口语文档检索领域的专门分支,口语命名实体识别因其研究有限及数据集匮乏而发展滞后。此外,跨语言迁移学习在口语命名实体识别中仍未得到探索。本文利用管道式和端到端方案,在荷兰语、英语和德语之间开展迁移学习研究。我们采用Wav2Vec2-XLS-R模型处理自定义伪标注数据集,并探究了多种架构以提升跨语言系统的适应性。实验结果表明,在标注数据有限的条件下,端到端口语命名实体识别表现优于基于管道式的替代方案。值得注意的是,从德语到荷兰语的迁移学习效果分别比荷兰语端到端系统和管道式系统高出7%和4%。本研究不仅验证了迁移学习在口语命名实体识别中的可行性,更展现出令人期待的未来评估前景,暗示着需要通过全面数据收集来进一步强化研究结果。