Selective inference is the problem of giving valid answers to statistical questions chosen in a data-driven manner. A standard solution to selective inference is simultaneous inference, which delivers valid answers to the set of all questions that could possibly have been asked. However, simultaneous inference can be unnecessarily conservative if this set includes many questions that were unlikely to be asked in the first place. We introduce a less conservative solution to selective inference that we call locally simultaneous inference, which only answers those questions that could plausibly have been asked in light of the observed data, all the while preserving rigorous type I error guarantees. For example, if the objective is to construct a confidence interval for the "winning" treatment effect in a clinical trial with multiple treatments, and it is obvious in hindsight that only one treatment had a chance to win, then our approach will return an interval that is nearly the same as the uncorrected, standard interval. Under mild conditions satisfied by common confidence intervals, locally simultaneous inference strictly dominates simultaneous inference, meaning it can never yield less statistical power but only more. Compared to conditional selective inference, which demands stronger guarantees, locally simultaneous inference is more easily applicable in nonparametric settings and is more numerically stable.
翻译:选择性推断是指对以数据驱动方式选择的统计问题提供有效答案的难题。同时推断作为标准解决方案,能为所有可能被提出的问题提供有效答案,但若问题集合包含大量原本不太可能被提出的问题,该方法可能过于保守。我们提出一种更不保守的选择性推断方法——局部同时推断,该方法仅回答基于观测数据具有现实可能性的问题,同时严格保证第一类错误率控制。例如,在包含多种处理的临床试验中构建"优胜"处理效应的置信区间时,若事后明显只有一种处理具有胜出可能,我们的方法将返回与未校正标准区间几乎一致的区间。在常见置信区间满足的温和条件下,局部同时推断严格优于同时推断,即绝不会降低统计功效而只会提升。相较于要求更强保证的条件选择性推断,局部同时推断更易应用于非参数场景且数值稳定性更强。