This paper is an extension of our previous conference paper. In recent years, there has been a growing interest among researchers in developing and improving speech recognition systems to facilitate and enhance human-computer interaction. Today, Automatic Speech Recognition (ASR) systems have become ubiquitous, used in everything from games to translation systems, robots, and more. However, much research is still needed on speech recognition systems for low-resource languages. This article focuses on the recognition of individual words in the Dari language using the Mel-frequency cepstral coefficients (MFCCs) feature extraction method and three different deep neural network models: Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), and Multilayer Perceptron (MLP), as well as two hybrid models combining CNN and RNN. We evaluate these models using an isolated Dari word corpus that we have created, consisting of 1000 utterances for 20 short Dari terms. Our study achieved an impressive average accuracy of 98.365%.
翻译:本文是我们在先前会议论文基础上的扩展研究。近年来,研究者们对开发和完善语音识别系统以促进和增强人机交互的兴趣日益增长。如今,自动语音识别(ASR)系统已无处不在,从游戏到翻译系统、机器人等领域均有应用。然而,针对低资源语言的语音识别系统仍亟需大量研究。本文聚焦于达里语(Dari language)的孤立词识别,采用梅尔频率倒谱系数(MFCCs)特征提取方法,并运用三种不同的深度神经网络模型:卷积神经网络(CNN)、循环神经网络(RNN)和多层感知机(MLP),以及两种结合CNN与RNN的混合模型。我们使用自建的达里语孤立词语料库(包含20个短达里语词汇的1000条语音样本)对上述模型进行评估。研究实现了令人瞩目的98.365%的平均准确率。