Deep learning-based ear disease diagnosis technology has proven effective and affordable. However, due to the lack of ear endoscope datasets with diversity, the practical potential of the deep learning model has not been thoroughly studied. Moreover, existing research failed to achieve a good trade-off between model inference speed and parameter size, rendering models inapplicable in real-world settings. To address these challenges, we constructed the first large-scale ear endoscopic dataset comprising eight types of ear diseases and disease-free samples from two institutions. Inspired by ShuffleNetV2, we proposed Best-EarNet, an ultrafast and ultralight network enabling real-time ear disease diagnosis. Best-EarNet incorporates a novel Local-Global Spatial Feature Fusion Module and multi-scale supervision strategy, which facilitates the model focusing on global-local information within feature maps at various levels. Utilizing transfer learning, the accuracy of Best-EarNet with only 0.77M parameters achieves 95.23% (internal 22,581 images) and 92.14% (external 1,652 images), respectively. In particular, it achieves an average frame per second of 80 on the CPU. From the perspective of model practicality, the proposed Best-EarNet is superior to state-of-the-art backbone models in ear lesion detection tasks. Most importantly, Ear-keeper, an intelligent diagnosis system based Best-EarNet, was developed successfully and deployed on common electronic devices (smartphone, tablet computer and personal computer). In the future, Ear-Keeper has the potential to assist the public and healthcare providers in performing comprehensive scanning and diagnosis of the ear canal in real-time video, thereby promptly detecting ear lesions.
翻译:基于深度学习的耳部疾病诊断技术已被证实有效且成本低廉。然而,由于缺乏具有多样性的耳内镜数据集,深度学习模型的实际潜力尚未得到充分研究。此外,现有研究未能实现模型推理速度与参数规模之间的良好平衡,导致模型在真实场景中难以应用。为应对这些挑战,我们构建了首个涵盖八种耳部疾病及来自两个机构的无疾病样本的大规模耳内镜数据集。受ShuffleNetV2启发,我们提出了Best-EarNet——一种实现实时耳部疾病诊断的超快超轻网络。Best-EarNet引入了新颖的局部-全局空间特征融合模块与多尺度监督策略,有助于模型在不同层级关注特征图中的全局-局部信息。通过迁移学习,仅含0.77M参数的Best-EarNet在内部数据集(22,581张图像)和外部数据集(1,652张图像)上的准确率分别达到95.23%和92.14%。特别地,其在CPU上实现了平均每秒80帧的处理速度。从模型实用性角度而言,所提出的Best-EarNet在耳部病变检测任务中优于现有最优骨干模型。最重要的是,基于Best-EarNet的智能诊断系统Ear-Keeper已成功开发并部署于常见电子设备(智能手机、平板电脑和个人计算机)。未来,Ear-Keeper有望辅助公众及医疗人员在实时视频中对耳道进行全面扫描与诊断,从而及时检测耳部病变。