Air pollution stands as the fourth leading cause of death globally. While extensive research has been conducted in this domain, most approaches rely on large datasets when it comes to prediction. This limits their applicability in low-resource settings though more vulnerable. This study addresses this gap by proposing a novel machine learning approach for accurate air quality prediction using two months of air quality data. By leveraging the World Weather Repository, the meteorological, air pollutant, and Air Quality Index features from 197 capital cities were considered to predict air quality for the next day. The evaluation of several machine learning models demonstrates the effectiveness of the Random Forest algorithm in generating reliable predictions, particularly when applied to classification rather than regression, approach which enhances the model's generalizability by 42%, achieving a cross-validation score of 0.38 for regression and 0.89 for classification. To instill confidence in the predictions, interpretable machine learning was considered. Finally, a cost estimation comparing the implementation of this solution in high-resource and low-resource settings is presented including a tentative of technology licensing business model. This research highlights the potential for resource-limited countries to independently predict air quality while awaiting larger datasets to further refine their predictions.
翻译:空气污染是全球第四大致死原因。尽管该领域已有大量研究,但大多数预测方法依赖大规模数据集,这限制了其在更易受影响的低资源环境中的适用性。本研究通过提出一种新颖的机器学习方法填补了这一空白,该方法仅使用两个月的空气质量数据即可实现精准预测。通过利用世界天气存储库,本研究考虑了197个首都城市的气象、空气污染物及空气质量指数特征,用于预测次日的空气质量。对多种机器学习模型的评估表明,随机森林算法在生成可靠预测方面表现优异,尤其适用于分类任务而非回归任务——这种方法将模型的泛化能力提升了42%,交叉验证得分分别为回归0.38和分类0.89。为增强预测的可信度,本研究采用了可解释机器学习技术。最后,本文给出了在高资源与低资源环境下实施该解决方案的成本估算,并提出了技术许可商业模式构想。这项研究凸显了资源有限国家在等待更大规模数据集以完善预测的同时,独立预测空气质量的潜在能力。