There are a variety of features of the human voice that can be classified as pitch, timbre, loudness, and vocal tone. It is observed in numerous incidents that human expresses their feelings using different vocal qualities when they are speaking. The primary objective of this research is to recognize different emotions of human beings such as anger, sadness, fear, neutrality, disgust, pleasant surprise, and happiness by using several MATLAB functions namely, spectral descriptors, periodicity, and harmonicity. To accomplish the work, we analyze the CREMA-D (Crowd-sourced Emotional Multimodal Actors Data) & TESS (Toronto Emotional Speech Set) datasets of human speech. The audio file contains data that have various characteristics (e.g., noisy, speedy, slow) thereby the efficiency of the ML (Machine Learning) models increases significantly. The EMD (Empirical Mode Decomposition) is utilized for the process of signal decomposition. Then, the features are extracted through the use of several techniques such as the MFCC, GTCC, spectral centroid, roll-off point, entropy, spread, flux, harmonic ratio, energy, skewness, flatness, and audio delta. The data is trained using some renowned ML models namely, Support Vector Machine, Neural Network, Ensemble, and KNN. The algorithms show an accuracy of 67.7%, 63.3%, 61.6%, and 59.0% respectively for the test data and 77.7%, 76.1%, 99.1%, and 61.2% for the training data. We have conducted experiments using Matlab and the result shows that our model is very prominent and flexible than existing similar works.
翻译:人类语音具有多种可分类的特征,如音高、音色、响度和语调。大量实例表明,人在说话时会通过不同的嗓音特质表达情感。本研究主要目标在于利用若干MATLAB函数(包括频谱描述符、周期性和谐波性)识别人类的不同情绪状态,如愤怒、悲伤、恐惧、中性、厌恶、惊喜及快乐。为此,我们分析了CREMA-D(众包情感多模态演员数据集)与TESS(多伦多情感语音集)两个人类语音数据集。音频文件包含多种特性(如嘈杂、快速、缓慢),这使得机器学习(ML)模型的效率显著提升。采用经验模态分解(EMD)进行信号分解处理,随后通过MFCC、GTCC、频谱质心、滚降点、熵、离散度、通量、谐波比、能量、偏度、平坦度及音频增量等多种技术提取特征。数据使用支持向量机、神经网络、集成学习及KNN等知名机器学习模型进行训练。测试数据显示各算法准确率分别为67.7%、63.3%、61.6%和59.0%,训练数据准确率则达77.7%、76.1%、99.1%及61.2%。我们在MATLAB环境中进行实验,结果表明本模型较现有同类工作具有显著优势与更强灵活性。