Automatic Speech Recognition (ASR) systems are a crucial technology that is used today to design a wide variety of applications, most notably, smart assistants, such as Alexa. ASR systems are essentially dialogue systems that employ Spoken Language Understanding (SLU) to extract meaningful information from speech. The main challenge with designing such systems is that they require a huge amount of labeled clean data to perform competitively, such data is extremely hard to collect and annotate to respective SLU tasks, furthermore, when designing such systems for low resource languages, where data is extremely limited, the severity of the problem intensifies. In this paper, we focus on a fairly popular SLU task, that is, Intent Classification while working with a low resource language, namely, Flemish. Intent Classification is a task concerned with understanding the intents of the user interacting with the system. We build on existing light models for intent classification in Flemish, and our main contribution is applying different augmentation techniques on two levels -- the voice level, and the phonetic transcripts level -- to the existing models to counter the problem of scarce labeled data in low-resource languages. We find that our data augmentation techniques, on both levels, have improved the model performance on a number of tasks.
翻译:自动语音识别(ASR)系统是当今设计各类应用——尤其是智能助手(如Alexa)——的关键技术。ASR系统本质上是采用口语语言理解(SLU)从语音中提取有意义信息的对话系统。设计此类系统的主要挑战在于:其需要海量标注清洁数据才能具备竞争性表现,而针对相应SLU任务收集并标注此类数据极为困难;更严重的是,当为数据极度匮乏的低资源语言设计此类系统时,问题难度将急剧增加。本文聚焦于一项较为流行的SLU任务——意图分类,并以低资源语言佛兰德语为研究对象。意图分类旨在理解与系统交互用户的意图。我们在现有轻量级佛兰德语意图分类模型基础上,主要贡献在于从语音层面和音标转录层面两个维度对现有模型应用不同的数据增强技术,以应对低资源语言中标注数据稀缺的问题。实验表明,我们提出的双层数据增强技术在多项任务上均提升了模型性能。