Spoken Language Understanding (SLU) aims to extract the semantic information from the speech utterance of user queries. It is a core component in a task-oriented dialogue system. With the spectacular progress of deep neural network models and the evolution of pre-trained language models, SLU has obtained significant breakthroughs. However, only a few high-resource languages have taken advantage of this progress due to the absence of SLU resources. In this paper, we seek to mitigate this obstacle by introducing SLURP-TN. This dataset was created by recording 55 native speakers uttering sentences in Tunisian dialect, manually translated from six SLURP domains. The result is an SLU Tunisian dialect dataset that comprises 4165 sentences recorded into around 5 hours of acoustic material. We also develop a number of Automatic Speech Recognition and SLU models exploiting SLUTP-TN. The Dataset and baseline models are available at: https://huggingface.co/datasets/Elyadata/SLURP-TN.
翻译:口语理解旨在从用户查询的语音中提取语义信息,是任务导向型对话系统的核心组件。得益于深度神经网络模型的显著进步和预训练语言模型的演进,口语理解已取得重大突破。然而,由于缺乏相关资源,仅有少数高资源语言能够充分受益于此进展。本文通过引入SLURP-TN数据集来缓解这一障碍。该数据集由55名母语者录制突尼斯方言句子而成,这些句子经过人工翻译自六个SLURP领域,最终形成了包含4165句语音(约5小时声学材料)的突尼斯方言口语理解数据集。我们还利用SLURP-TN开发了多个自动语音识别与口语理解基线模型。数据集及基线模型已开源在:https://huggingface.co/datasets/Elyadata/SLURP-TN。