Most scientific disciplines use significance testing to draw conclusions about experimental or observational data. This classical approach provides a theoretical guarantee for controlling the number of false positives across a set of hypothesis tests, making it an appealing framework for scientists seeking to limit the number of false effects or associations that they claim to observe. Unfortunately, this theoretical guarantee applies to few experiments, and the true false positive rate (FPR) is much higher. Scientists have plenty of freedom to choose the error rate to control, the tests to include in the adjustment, and the method of correction, making strong error control difficult to attain. In addition, hypotheses are often tested after finding unexpected relationships or patterns, the data are analysed in several ways, and analyses may be run repeatedly as data accumulate. As a result, adjusted p-values are too small, incorrect conclusions are often reached, and results are harder to reproduce. In the following, I argue why the FPR is rarely controlled meaningfully and why shrinking parameter estimates is preferable to p-value adjustments.
翻译:大多数科学领域在分析实验或观测数据时,都采用显著性检验方法得出推论。这一经典方法为控制一系列假设检验中的假阳性数量提供了理论保障,因此成为科学家们用以限制声称观察到的虚假效应或关联的有力工具。遗憾的是,这种理论保障仅适用于少数实验,而真实的假阳性率(FPR)要高得多。科学家们在选择要控制的错误率、纳入校正的检验以及校正方法方面拥有相当大的自由度,这使得强效的错误控制难以实现。此外,假设往往在发现意外关系或模式后才被检验,数据被以多种方式分析,且分析过程可能随着数据积累而重复进行。结果导致校正后的p值过小,经常得出错误结论,且研究结果更难以被复现。下文将论证为何FPR很少能得到有意义的控制,以及为何收缩参数估计值比调整p值更为可取。