A key strategy in societal adaptation to climate change is the use of alert systems to reduce the adverse health impacts of extreme heat events by prompting preventative action. In this work, we investigate reinforcement learning (RL) as a tool to optimize the effectiveness of such systems. Our contributions are threefold. First, we introduce a novel RL environment enabling the evaluation of the effectiveness of heat alert policies to reduce heat-related hospitalizations. The rewards model is trained from a comprehensive dataset of historical weather, Medicare health records, and socioeconomic/geographic features. We use variational Bayesian techniques to address low-signal effects and spatial heterogeneity, which are commonly encountered in climate & health settings. The transition model incorporates real historical weather patterns enriched by a data augmentation mechanism based on climate region similarity. Second, we use this environment to evaluate standard RL algorithms in the context of heat alert issuance. Our analysis shows that policy constraints are needed to improve the initially poor performance of RL. Lastly, a post hoc contrastive analysis provides insight into scenarios where our modified heat alert-RL policies yield significant gains/losses over the current National Weather Service alert policy in the United States.
翻译:在应对气候变化的社会适应策略中,预警系统通过促使人们采取预防性行动来减少极端高温事件对健康的不利影响。本研究探讨了强化学习作为优化此类系统效能的工具。我们的贡献包含三个方面:首先,我们构建了一个新的强化学习环境,用于评估高温预警政策在减少热相关住院方面的有效性。奖励模型基于历史天气、医疗保险健康记录及社会经济/地理特征的综合数据集进行训练,采用变分贝叶斯技术处理气候与健康场景中常见的低信号效应和空间异质性。转移模型融合了实际历史天气模式,并通过基于气候区域相似性的数据增强机制进行扩展。其次,我们利用该环境评估了标准强化学习算法在高温预警发布中的表现。分析表明,需要施加政策约束来改善强化学习初始较差的性能。最后,事后对比分析揭示了在特定场景下,我们改进的高温预警-强化学习政策相较于美国国家气象局现行预警政策产生的显著收益/损失。