To mitigate the impact of the pandemic, several measures include lockdowns, rapid vaccination programs, school closures, and economic stimulus. These interventions can have positive or unintended negative consequences. Current research to model and determine an optimal intervention automatically through round-tripping is limited by the simulation objectives, scale (a few thousand individuals), model types that are not suited for intervention studies, and the number of intervention strategies they can explore (discrete vs continuous). We address these challenges using a Deep Deterministic Policy Gradient (DDPG) based policy optimization framework on a large-scale (100,000 individual) epidemiological agent-based simulation where we perform multi-objective optimization. We determine the optimal policy for lockdown and vaccination in a minimalist age-stratified multi-vaccine scenario with a basic simulation for economic activity. With no lockdown and vaccination (mid-age and elderly), results show optimal economy (individuals below the poverty line) with balanced health objectives (infection, and hospitalization). An in-depth simulation is needed to further validate our results and open-source our framework.
翻译:为减轻大流行的影响,多种措施包括封锁、快速疫苗接种计划、学校关闭和经济刺激。这些干预措施可能产生积极或意想不到的负面后果。当前通过迭代优化自动建模并确定最优干预策略的研究受到模拟目标、规模(数千个体)、不适用于干预研究的模型类型以及其可探索的干预策略数量(离散与连续)的限制。我们利用基于深度确定性策略梯度(DDPG)的策略优化框架,在大型(10万个个体)流行病学智能体模拟中解决这些挑战,并执行多目标优化。我们在一个包含经济活动的简化年龄分层多疫苗场景中确定封锁和疫苗接种的最优策略。无封锁和疫苗接种(中年和老年人)的结果显示,在平衡健康目标(感染和住院)下,经济表现最优(个体处于贫困线以下)。需要更深入的模拟来进一步验证我们的结果,并将我们的框架开源。