Post-selection inference has recently been proposed as a way of quantifying uncertainty about detected changepoints. The idea is to run a changepoint detection algorithm, and then re-use the same data to perform a test for a change near each of the detected changes. By defining the p-value for the test appropriately, so that it is conditional on the information used to choose the test, this approach will produce valid p-values. We show how to improve the power of these procedures by conditioning on less information. This gives rise to an ideal selective p-value that is intractable but can be approximated by Monte Carlo. We show that for any Monte Carlo sample size, this procedure produces valid p-values, and empirically that noticeable increase in power is possible with only very modest Monte Carlo sample sizes. Our procedure is easy to implement given existing post-selection inference methods, as we just need to generate perturbations of the data set and re-apply the post-selection method to each of these. On genomic data consisting of human GC content, our procedure increases the number of significant changepoints that are detected from e.g. 17 to 27, when compared to existing methods.
翻译:后选择推断近年来被提出作为一种量化检测变点不确定性的方法。其核心思路是:先运行变点检测算法,然后利用同一数据在检测到的每个变点附近检验是否存在变化。通过合理定义检验的p值,使其在依赖用于选择检验的信息条件下保持有效,该方法能生成有效的p值。我们展示了如何通过减少条件化信息来提升这些程序的检验功效。这催生了一种理想的选择性p值,虽然计算困难,但可通过蒙特卡洛方法近似求解。我们证明:对于任意蒙特卡洛样本量,该程序均能生成有效p值,且实验表明,即使使用非常小的蒙特卡洛样本量,也能实现显著的检验功效提升。基于现有后选择推断方法,我们的程序易于实现——仅需生成数据集的扰动版本并对每个扰动重新应用后选择方法。在包含人类GC含量的基因组数据上,与传统方法相比,我们的程序将检测到的显著变点数量从17个提升至27个。