Skin lesion classification is essential for early dermatological diagnosis, yet many existing computer-aided systems rely primarily on dermoscopic images and underutilize the multimodal evidence routinely available in clinical practice. To address this gap, we propose \textbf{JI-ADF}, a trimodal deep learning framework that integrates dermoscopic images, clinical photographs, and structured patient metadata for clinically grounded skin lesion classification. The proposed architecture combines joint multimodal representation learning with modality-specific auxiliary supervision and an adaptive decision fusion mechanism that dynamically calibrates modality contributions on a per-sample basis. To enhance cross-modal reasoning while preserving modality-specific evidence, we further introduce a multimodal fusion attention (MMFA) module. We evaluate JI-ADF on the large-scale MILK10k benchmark, which reflects real-world clinical acquisition conditions and severe class imbalance. The proposed method demonstrates strong and well-balanced performance across lesion categories, improving sensitivity and Dice score while maintaining high specificity and good calibration. Extensive analyses, including modality ablation, calibration evaluation, and Grad-CAM visualization, further confirm the robustness and clinically meaningful behavior of the model. These results indicate that JI-ADF provides a reliable and practical foundation for multimodal skin lesion classification in real-world clinical settings.
翻译:摘要:皮肤病变分类对于早期皮肤病诊断至关重要,然而许多现有计算机辅助系统主要依赖皮肤镜图像,未能充分利用临床实践中常规可获得的多模态证据。为弥补这一不足,我们提出JI-ADF——一种三模态深度学习框架,该框架整合了皮肤镜图像、临床照片及结构化患者元数据,用于具有临床依据的皮肤病变分类。所提出的架构将联合多模态表示学习与模态特定辅助监督相结合,并引入自适应决策融合机制,可在逐样本基础上动态校准各模态的贡献。为进一步增强跨模态推理能力同时保留模态特异性证据,我们引入了多模态融合注意力(MMFA)模块。我们在反映真实临床采集条件及严重类别不平衡的大规模MILK10k基准上评估了JI-ADF。该方法在各类病变上均展现出稳健且均衡的性能,在保持高特异性和良好校准能力的同时,提升了敏感性与Dice分数。通过消融实验、校准评估及Grad-CAM可视化等广泛分析,进一步证实了模型的鲁棒性及临床意义。上述结果表明,JI-ADF为真实临床环境下的多模态皮肤病变分类提供了可靠且实用的基础。