In this paper, SER_AMPEL, a multi-source dataset for speech emotion recognition (SER) is presented. The peculiarity of the dataset is that it is collected with the aim of providing a reference for speech emotion recognition in case of Italian older adults. The dataset is collected following different protocols, in particular considering acted conversations, extracted from movies and TV series, and recording natural conversations where the emotions are elicited by proper questions. The evidence of the need for such a dataset emerges from the analysis of the state of the art. Preliminary considerations on the critical issues of SER are reported analyzing the classification results on a subset of the proposed dataset.
翻译:本文介绍了一个用于语音情感识别(SER)的多源数据集SER_AMPEL。该数据集的特殊之处在于,其收集目的旨在为意大利老年人语音情感识别提供参考基准。数据集遵循不同协议进行采集,具体包括从电影和电视剧中提取的对话表演,以及通过适当问题引发情感的自然对话录音。对现有技术现状的分析揭示了此类数据集的必要性。通过分析所提出数据集子集上的分类结果,本文报告了对语音情感识别关键问题的初步思考。