Selective inference is the problem of giving valid answers to statistical questions chosen in a data-driven manner. A standard solution to selective inference is simultaneous inference, which delivers valid answers to the set of all questions that could possibly have been asked. However, simultaneous inference can be unnecessarily conservative if this set includes many questions that were unlikely to be asked in the first place. We introduce a less conservative solution to selective inference that we call locally simultaneous inference, which only answers those questions that could plausibly have been asked in light of the observed data, all the while preserving rigorous type I error guarantees. For example, if the objective is to construct a confidence interval for the "winning" treatment effect in a clinical trial with multiple treatments, and it is obvious in hindsight that only one treatment had a chance to win, then our approach will return an interval that is nearly the same as the uncorrected, standard interval. Compared to conditional selective inference, which demands stronger, conditional guarantees, locally simultaneous inference is more easily applicable in nonparametric settings and is more numerically stable.
翻译:选择性推断是一个在数据驱动方式下对所选统计问题提供有效答案的难题。同时推断作为标准解决方案,能为所有可能被提出的问题集合提供有效答案。然而,当这个集合包含许多最初几乎不可能被提出的问题时,同时推断可能过于保守。我们提出一种更少保守性的选择性推断方法,称为局部同时推断,该方法仅回答基于观测数据看似合理的问题,同时严格保证第一类错误率控制。例如,若目标是为多治疗组临床试验中的"获胜"治疗效果构建置信区间,且事后明显只有一种治疗有可能胜出,我们的方法将返回一个几乎与未校正的标准区间相同的区间。与要求更强条件性保证的条件选择性推断相比,局部同时推断在非参数环境中更易应用且数值稳定性更高。