Recent work has reported that AI classifiers trained on audio recordings can accurately predict severe acute respiratory syndrome coronavirus 2 (SARSCoV2) infection status. Here, we undertake a large scale study of audio-based deep learning classifiers, as part of the UK governments pandemic response. We collect and analyse a dataset of audio recordings from 67,842 individuals with linked metadata, including reverse transcription polymerase chain reaction (PCR) test outcomes, of whom 23,514 tested positive for SARS CoV 2. Subjects were recruited via the UK governments National Health Service Test-and-Trace programme and the REal-time Assessment of Community Transmission (REACT) randomised surveillance survey. In an unadjusted analysis of our dataset AI classifiers predict SARS-CoV-2 infection status with high accuracy (Receiver Operating Characteristic Area Under the Curve (ROCAUC) 0.846 [0.838, 0.854]) consistent with the findings of previous studies. However, after matching on measured confounders, such as age, gender, and self reported symptoms, our classifiers performance is much weaker (ROC-AUC 0.619 [0.594, 0.644]). Upon quantifying the utility of audio based classifiers in practical settings, we find them to be outperformed by simple predictive scores based on user reported symptoms.
翻译:近期研究表明,基于音频数据训练的AI分类器能够准确预测严重急性呼吸综合征冠状病毒2型(SARS-CoV-2)感染状态。本研究作为英国政府大流行病应对措施的一部分,开展了一项大规模音频深度学习分类器研究。我们收集并分析了来自67,842名个体的音频记录数据集,并关联其元数据(包括逆转录聚合酶链式反应(PCR)检测结果),其中23,514人检测结果呈SARS-CoV-2阳性。受试者通过英国政府国民健康服务检测与追踪计划以及社区传播实时评估(REACT)随机监测调查招募。在对数据集进行未经调整的分析中,AI分类器以高准确率预测SARS-CoV-2感染状态(受试者工作特征曲线下面积(ROC-AUC)为0.846[0.838, 0.854]),与先前研究结果一致。然而,在根据年龄、性别和自我报告症状等已测量混杂因素进行匹配后,分类器性能显著下降(ROC-AUC为0.619[0.594, 0.644])。通过量化基于音频分类器在实际场景中的效用,我们发现其表现劣于基于用户报告症状的简单预测评分。