Large-scale administrative or observational datasets are increasingly used to inform decision making. While this effort aims to ground policy in real-world evidence, challenges have arise as that selection bias and other forms of distribution shift often plague observational data. Previous attempts to provide robust inferences have given guarantees depending on a user-specified amount of possible distribution shift (e.g., the maximum KL divergence between the observed and target distributions). However, decision makers will often have additional knowledge about the target distribution which constrains the kind of shifts which are possible. To leverage such information, we proposed a framework that enables statistical inference in the presence of distribution shifts which obey user-specified constraints in the form of functions whose expectation is known under the target distribution. The output is high-probability bounds on the value an estimand takes on the target distribution. Hence, our method leverages domain knowledge in order to partially identify a wide class of estimands. We analyze the computational and statistical properties of methods to estimate these bounds, and show that our method can produce informative bounds on a variety of simulated and semisynthetic tasks.
翻译:大规模行政或观测数据集正日益被用于辅助决策制定。尽管此举旨在将政策建立在真实世界证据基础上,但选择偏差及其他形式的分布偏移常困扰观测数据,因而引发诸多挑战。以往提供稳健推断的方法依赖于用户指定的可能分布偏移量(例如观测分布与目标分布之间的最大KL散度)来给出保证。然而,决策者往往掌握关于目标分布的额外知识,从而限制了可能发生的偏移类型。为利用此类信息,我们提出了一个框架,该框架能在遵循用户指定约束(以函数形式呈现,且这些函数在目标分布下的期望值已知)的分布偏移条件下进行统计推断。其输出是目标分布上待估参数值的高概率界值。因此,我们的方法可借助领域知识对广泛类别的待估参数进行部分识别。我们分析了估算这些界值方法的计算与统计特性,并表明该方法能在多种模拟及半合成任务中生成有信息量的界值。