Respiratory diseases remain a leading cause of mortality worldwide, highlighting the need for faster and more accurate diagnostic tools. This work presents a novel approach leveraging digital stethoscope technology for automatic respiratory disease classification and biometric analysis. Our approach has the potential to significantly enhance traditional auscultation practices. By leveraging one of the largest publicly available medical database of respiratory sounds, we train machine learning models to classify various respiratory health conditions. Our method differs from conventional methods by using Empirical Mode Decomposition (EMD) and spectral analysis techniques to isolate clinically relevant biosignals embedded within acoustic data captured by digital stethoscopes. This approach focuses on information closely tied to cardiovascular and respiratory patterns within the acoustic data. Spectral analysis and filtering techniques isolate Intrinsic Mode Functions (IMFs) strongly correlated with these physiological phenomena. These biosignals undergo a comprehensive feature extraction process for predictive modeling. These features then serve as input to train several machine learning models for both classification and regression tasks. Our approach achieves high accuracy in both binary classification (89% balanced accuracy for healthy vs. diseased) and multi-class classification (72% balanced accuracy for specific diseases like pneumonia and COPD). For the first time, this work introduces regression models capable of estimating age and body mass index (BMI) based solely on acoustic data, as well as a model for sex classification. Our findings underscore the potential of intelligent digital stethoscopes to significantly enhance assistive and remote diagnostic capabilities, contributing to advancements in digital health, telehealth, and remote patient monitoring.
翻译:呼吸系统疾病仍然是全球范围内导致死亡的主要原因之一,凸显了对更快速、更精准诊断工具的迫切需求。本研究提出了一种创新方法,利用数字听诊器技术实现呼吸系统疾病的自动分类与生物特征分析,该方法有望显著提升传统听诊实践。通过利用目前规模最大的公开呼吸音医学数据库之一,我们训练机器学习模型对多种呼吸系统健康状况进行分类。与常规方法不同,我们的研究采用经验模态分解(EMD)和频谱分析技术,从数字听诊器捕获的声学数据中分离出具有临床关联性的生物信号。该方法聚焦于声学数据中与心血管和呼吸模式紧密相关的信息。通过频谱分析与滤波技术,我们提取出与这些生理现象高度相关的本征模态函数(IMF)。这些生物信号经过全面的特征提取流程后,用于构建预测模型。上述特征被输入至多个机器学习模型,以完成分类与回归任务。我们的方法在二分类任务中实现了高准确率(健康与患病状态的平衡准确率为89%),在多分类任务中亦表现优异(肺炎、慢阻肺等特定疾病的平衡准确率为72%)。本研究首次引入了仅基于声学数据估算年龄与身体质量指数(BMI)的回归模型,以及性别分类模型。研究结果表明,智能数字听诊器在显著增强辅助诊断与远程诊断能力方面具有巨大潜力,将为数字健康、远程医疗及远程患者监护等领域的发展做出贡献。