A main research goal in various studies is to use an observational data set and provide a new set of counterfactual guidelines that can yield causal improvements. Dynamic Treatment Regimes (DTRs) are widely studied to formalize this process. However, available methods in finding optimal DTRs often rely on assumptions that are violated in real-world applications (e.g., medical decision-making or public policy), especially when (a) the existence of unobserved confounders cannot be ignored, and (b) the unobserved confounders are time-varying (e.g., affected by previous actions). When such assumptions are violated, one often faces ambiguity regarding the underlying causal model. This ambiguity is inevitable, since the dynamics of unobserved confounders and their causal impact on the observed part of the data cannot be understood from the observed data. Motivated by a case study of finding superior treatment regimes for patients who underwent transplantation in our partner hospital and faced a medical condition known as New Onset Diabetes After Transplantation (NODAT), we extend DTRs to a new class termed Ambiguous Dynamic Treatment Regimes (ADTRs), in which the causal impact of treatment regimes is evaluated based on a "cloud" of causal models. We then connect ADTRs to Ambiguous Partially Observable Mark Decision Processes (APOMDPs) and develop Reinforcement Learning methods, which enable using the observed data to efficiently learn an optimal treatment regime. We establish theoretical results for these learning methods, including (weak) consistency and asymptotic normality. We further evaluate the performance of these learning methods both in our case study and in simulation experiments.
翻译:各类研究的主要目标之一是使用观察性数据集,提供一套能产生因果改进的反事实指南。动态治疗方案(DTRs)被广泛研究以形式化这一过程。然而,现有寻找最优DTRs的方法通常依赖于在实际应用(如医疗决策或公共政策)中可能被违反的假设,特别是当(a)未观测混杂因素的存在不可忽略,且(b)未观测混杂因素随时间变化(例如受先前行动影响)时。当这些假设被违反时,人们常常面临关于潜在因果模型的不确定性。这种不确定性是不可避免的,因为未观测混杂因素的动态及其对数据观测部分的因果影响无法从观测数据中推断。受一项案例研究的启发——该案例旨在为我院移植术后发生移植后新发糖尿病(NODAT)的患者寻找更优的治疗方案,我们将DTRs扩展为一类新方法,称为模糊动态治疗方案(ADTRs),其中治疗方案因果效应的评估基于一组因果模型的“云集”。随后,我们将ADTRs与模糊部分可观测马尔可夫决策过程(APOMDPs)联系起来,并开发了强化学习方法,从而能够利用观测数据高效学习最优治疗方案。我们为这些学习方法建立了理论结果,包括(弱)相合性和渐近正态性。此外,我们还在案例研究和模拟实验中评估了这些学习方法的性能。