Decision-makers often deploy the best-performing treatment from a randomized experiment, creating a winner's curse: selection favors treatments whose observed outcomes are high partly because of statistical noise, so the naïve estimate of the winner is upward biased. We distinguish two forms of winner's curse, bias relative to the true best treatment (global) and bias relative to the selected treatment's true mean (selective), and link them to regret from deploying a suboptimal treatment. This framework defines seven decision-relevant evaluation targets: mean bias, mean squared error, and confidence interval coverage for the global and selective winner's curse, and mean regret. We then show that methods that perform well on one target can perform poorly on others, so corrections should be matched to the manager's objective. Across simulations with varying effect sizes, multiple-arm settings, and data calibrated to an online A/B testing platform, no method dominates uniformly: the plug-in estimator performs best when treatment differences are large, cross-fitting performs best when treatments are similar, and resampling methods often achieve low mean squared error for moderate differences. We also introduce an adaptive empirical likelihood procedure that delivers asymptotically valid confidence intervals across settings without the tuning sensitivity of resampling-based methods.
翻译:决策者通常从随机实验中部署表现最佳的处理方案,这会导致"胜出者诅咒":选择倾向于那些观测结果较高的处理方案——部分原因是统计噪声,因此对胜出者的朴素估计存在向上偏差。我们区分了两种形式的胜出者诅咒:相对于真实最佳处理方案的偏差(全局型)和相对于所选处理方案真实均值的偏差(选择型),并将它们与部署次优处理方案导致的遗憾相关联。该框架定义了七个与决策相关的评估目标:全局型和选择型胜出者诅咒的均值偏差、均方误差、置信区间覆盖率,以及均值遗憾。我们进一步表明,在某个目标上表现良好的方法可能在其他目标上表现欠佳,因此校正方法需与管理者目标相匹配。在包含不同效应量、多臂设置以及基于在线A/B测试平台校准数据的模拟实验中,没有任何方法能完全占优:当处理效应差异较大时,插件估计器表现最佳;当处理效应相似时,交叉拟合方法表现最优;而重采样方法在中等差异情况下通常能实现较低的均方误差。我们还引入了一种自适应经验似然方法,无需重采样方法对调参的敏感性,即可在不同设置下给出渐近有效的置信区间。