This paper presents the Coswara dataset, a dataset containing diverse set of respiratory sounds and rich meta-data, recorded between April-2020 and February-2022 from 2635 individuals (1819 SARS-CoV-2 negative, 674 positive, and 142 recovered subjects). The respiratory sounds contained nine sound categories associated with variants of breathing, cough and speech. The rich metadata contained demographic information associated with age, gender and geographic location, as well as the health information relating to the symptoms, pre-existing respiratory ailments, comorbidity and SARS-CoV-2 test status. Our study is the first of its kind to manually annotate the audio quality of the entire dataset (amounting to 65~hours) through manual listening. The paper summarizes the data collection procedure, demographic, symptoms and audio data information. A COVID-19 classifier based on bi-directional long short-term (BLSTM) architecture, is trained and evaluated on the different population sub-groups contained in the dataset to understand the bias/fairness of the model. This enabled the analysis of the impact of gender, geographic location, date of recording, and language proficiency on the COVID-19 detection performance.
翻译:本文介绍Coswara数据集,该数据集包含多样化的呼吸音及丰富的元数据,录制于2020年4月至2022年2月期间,涵盖2635名受试者(其中1819名SARS-CoV-2阴性、674名阳性、142名康复者)。呼吸音包含九类与呼吸、咳嗽及语言变体相关的音频类别。丰富的元数据涉及年龄、性别、地理位置的统计信息,以及与症状、既往呼吸系统疾病、合并症及SARS-CoV-2检测状态相关的健康信息。本研究首次通过人工监听方式对整组数据集(总计约65小时)进行音频质量标注。本文总结了数据采集流程、人口统计特征、症状及音频数据信息。我们基于双向长短期记忆架构训练并评估了COVID-19分类器,针对数据集中不同人群子集的分类表现进行测试,以理解模型的偏差/公平性。这一过程实现了对性别、地理位置、录音日期及语言熟练程度等因素影响COVID-19检测性能的分析。