The simultaneous estimation of many parameters based on data collected from corresponding studies is a key research problem that has received renewed attention in the high-dimensional setting. Many practical situations involve heterogeneous data where heterogeneity is captured by a nuisance parameter. Effectively pooling information across samples while correctly accounting for heterogeneity presents a significant challenge in large-scale estimation problems. We address this issue by introducing the ``Nonparametric Empirical Bayes Structural Tweedie" (NEST) estimator, which efficiently estimates the unknown effect sizes and properly adjusts for heterogeneity via a generalized version of Tweedie's formula. For the normal means problem, NEST simultaneously handles the two main selection biases introduced by heterogeneity: one, the selection bias in the mean, which cannot be effectively corrected without also correcting for, two, selection bias in the variance. We develop theory to show that NEST is asymptotically as good as the optimal Bayes rule that uniquely minimizes a weighted squared error loss. In our simulation studies NEST outperforms competing methods, with much efficiency gains in many settings. The proposed method is demonstrated on estimating the batting averages of baseball players and Sharpe ratios of mutual fund returns. Extensions to other members of the two-parameter exponential family are discussed.
翻译:基于对应研究收集的数据同时估计多个参数是一个关键研究问题,在高维场景下重新受到关注。许多实际情境涉及异质数据,其中异质性由干扰参数刻画。在大型估计问题中,有效整合跨样本信息并正确解释异质性构成了显著挑战。我们通过引入“非参数经验贝叶斯结构特威迪”(NEST)估计量来解决该问题,该估计量通过特威迪公式的广义版本高效估计未知效应量,并适当调整异质性。对于正态均值问题,NEST同时处理由异质性引入的两种主要选择偏差:一种是均值的选择偏差,若不校正第二种——方差的选择偏差,则无法有效修正。我们发展理论证明NEST在渐近意义上与最优贝叶斯规则同样有效,该规则唯一最小化加权平方误差损失。在模拟研究中,NEST优于竞争方法,并在许多设置中实现大幅效率提升。该方法通过估计棒球运动员击球率和共同基金回报夏普比率得到验证,并讨论了向双参数指数族其他成员的扩展。