This study investigates the interplay among social demographics, built environment characteristics, and environmental hazard exposure features in determining community level cancer prevalence. Utilizing data from five Metropolitan Statistical Areas in the United States: Chicago, Dallas, Houston, Los Angeles, and New York, the study implemented an XGBoost machine learning model to predict the extent of cancer prevalence and evaluate the importance of different features. Our model demonstrates reliable performance, with results indicating that age, minority status, and population density are among the most influential factors in cancer prevalence. We further explore urban development and design strategies that could mitigate cancer prevalence, focusing on green space, developed areas, and total emissions. Through a series of experimental evaluations based on causal inference, the results show that increasing green space and reducing developed areas and total emissions could alleviate cancer prevalence. The study and findings contribute to a better understanding of the interplay among urban features and community health and also show the value of interpretable machine learning models for integrated urban design to promote public health. The findings also provide actionable insights for urban planning and design, emphasizing the need for a multifaceted approach to addressing urban health disparities through integrated urban design strategies.
翻译:本研究探究了社会人口统计、建成环境特征与环境危害暴露因素在社区层面癌症患病率中的交互作用。利用美国五个大都市统计区(芝加哥、达拉斯、休斯顿、洛杉矶和纽约)的数据,本研究采用XGBoost机器学习模型预测癌症患病程度并评估不同特征的重要性。我们的模型表现出可靠性能,结果显示年龄、少数族裔状况和人口密度是影响癌症患病率的最关键因素。我们进一步探索了可能降低癌症患病率的城市发展与设计策略,重点关注绿地空间、开发区域面积和总排放量。通过一系列基于因果推断的实验评估,结果表明增加绿地空间、减少开发区域面积和总排放量可缓解癌症患病率。本研究及其发现有助于更深入理解城市特征与社区健康之间的相互作用机制,同时展示了可解释机器学习模型在促进公共健康的综合性城市设计中的价值。研究成果还为城市规划和设计提供了可操作见解,强调通过整合性城市设计策略来应对城市健康差异需要采取多方面综合方法。