Detecting Type 2 Diabetes (T2D) and Prediabetes (PD) is a real challenge for medicine due to the absence of pathogenic symptoms and the lack of known associated risk factors. Even though some proposals for machine learning models enable the identification of people at risk, the nature of the condition makes it so that a model suitable for one population may not necessarily be suitable for another. In this article, the development and assessment of predictive models to identify people at risk for T2D and PD specifically in Argentina are discussed. First, the database was thoroughly preprocessed and three specific datasets were generated considering a compromise between the number of records and the amount of available variables. After applying 5 different classification models, the results obtained show that a very good performance was observed for two datasets with some of these models. In particular, RF, DT, and ANN demonstrated great classification power, with good values for the metrics under consideration. Given the lack of this type of tool in Argentina, this work represents the first step towards the development of more sophisticated models.
翻译:由于缺乏致病症状及已知相关风险因素,对2型糖尿病(T2D)和糖尿病前期(PD)的检测仍是医学领域面临的真正挑战。尽管现有部分机器学习模型能够识别风险人群,但该疾病的性质决定了适用于某一群体的模型未必适用于另一群体。本文探讨了专门针对阿根廷地区开发的T2D及PD风险人群识别预测模型的构建与评估过程。首先对数据库进行了全面预处理,通过权衡记录数量与可用变量数目,生成了三个特定数据集。在应用五种不同分类模型后,结果表明其中两个数据集在部分模型上表现优异。特别值得注意的是,随机森林(RF)、决策树(DT)和人工神经网络(ANN)展现出强大的分类能力,各项评估指标均达到理想值。鉴于阿根廷缺乏此类工具,本研究为开发更精密模型奠定了基础。