Modern genomic studies are increasingly focused on discovering more and more interesting genes associated with a health response. Traditional shrinkage priors are primarily designed to detect a handful of signals from tens of thousands of predictors in the so-called ultra-sparsity domain. However, they may fail to identify signals when the degree of sparsity is moderate. Robust sparse estimation under diverse sparsity regimes relies on a tail-adaptive shrinkage property. In this property, the tail-heaviness of the prior adjusts adaptively, becoming larger or smaller as the sparsity level increases or decreases, respectively, to accommodate more or fewer signals. In this study, we propose a global-local-tail (GLT) Gaussian mixture distribution that ensures this property. We examine the role of the tail-index of the prior in relation to the underlying sparsity level and demonstrate that the GLT posterior contracts at the minimax optimal rate for sparse normal mean models. We apply both the GLT prior and the Horseshoe prior to real data problems and simulation examples. Our findings indicate that the varying tail rule based on the GLT prior offers advantages over a fixed tail rule based on the Horseshoe prior in diverse sparsity regimes.
翻译:现代基因组研究日益关注发现与健康反应相关的更多有趣基因。传统收缩先验主要设计用于在所谓的超稀疏域中从数万个预测变量中检测少量信号,但在稀疏程度适中时可能无法识别信号。不同稀疏性条件下的稳健稀疏估计依赖于尾自适应收缩特性——该特性中先验分布的尾部重度会自适应调整,随稀疏程度升高而增大、降低而减小,以适应更多或更少的信号。本研究提出一种确保该特性的全局-局部-尾(GLT)高斯混合分布。我们检验了先验尾部指数与潜在稀疏水平之间的关系,并证明在稀疏正态均值模型中,GLT后验以极小化极大最优速率收缩。我们将GLT先验和Horseshoe先验应用于实际数据问题及模拟案例,研究结果表明基于GLT先验的动态尾部规则在不同稀疏性条件下较基于Horseshoe先验的固定尾部规则更具优势。