Multiple testing is an important research direction that has gained major attention in recent years. Currently, most multiple testing procedures are designed with p-values or Local false discovery rate (Lfdr) statistics. However, p-values obtained by applying probability integral transform to some well-known test statistics often do not incorporate information from the alternatives, resulting in suboptimal procedures. On the other hand, Lfdr based procedures can be asymptotically optimal but their guarantee on false discovery rate (FDR) control relies on consistent estimation of Lfdr, which is often difficult in practice especially when the incorporation of side information is desirable. In this article, we propose a novel and flexibly constructed class of statistics, called rho-values, which combines the merits of both p-values and Lfdr while enjoys superiorities over methods based on these two types of statistics. Specifically, it unifies these two frameworks and operates in two steps, ranking and thresholding. The ranking produced by rho-values mimics that produced by Lfdr statistics, and the strategy for choosing the threshold is similar to that of p-value based procedures. Therefore, the proposed framework guarantees FDR control under weak assumptions; it maintains the integrity of the structural information encoded by the summary statistics and the auxiliary covariates and hence can be asymptotically optimal. We demonstrate the efficacy of the new framework through extensive simulations and two data applications.
翻译:多重检验是近年来备受关注的重要研究方向。目前,大多数多重检验方法基于p值或局部错误发现率(Local false discovery rate, Lfdr)统计量进行设计。然而,通过对某些著名检验统计量应用概率积分变换得到的p值,往往未能整合备择假设信息,导致方法非最优。另一方面,基于Lfdr的方法虽能渐近最优,但其对错误发现率(False discovery rate, FDR)控制的保障依赖于Lfdr的一致估计,在实际应用中尤其当需要整合辅助信息时,这一估计通常存在困难。本文提出一类新颖且可灵活构造的统计量——rho值,它融合了p值与Lfdr的优点,同时优于基于这两类统计量的方法。具体而言,该统计量统一了上述两种框架,并通过排序与阈值选择两步操作实现功能:由rho值产生的排序结果与Lfdr统计量产生的排序相似,而阈值选取策略则类似基于p值的方法。因此,该新框架能在弱假设下保证FDR控制,同时保持汇总统计量与辅助协变量所编码结构信息的完整性,从而具有渐近最优性。我们通过大量模拟实验和两项数据应用验证了该框架的有效性。