During multiple testing, researchers often adjust their alpha level to control the familywise error rate for a statistical inference about a joint union alternative hypothesis (e.g., "H1,1 or H1,2"). However, in some cases, they do not make this inference. Instead, they make separate inferences about each of the individual hypotheses that comprise the joint hypothesis (e.g., H1,1 and H1,2). For example, a researcher might use a Bonferroni correction to adjust their alpha level from the conventional level of 0.050 to 0.025 when testing H1,1 and H1,2, find a significant result for H1,1 (p < 0.025) and not for H1,2 (p > .0.025), and so claim support for H1,1 and not for H1,2. However, these separate individual inferences do not require an alpha adjustment. Only a statistical inference about the union alternative hypothesis "H1,1 or H1,2" requires an alpha adjustment because it is based on "at least one" significant result among the two tests, and so it refers to the familywise error rate. Hence, an inconsistent correction occurs when a researcher corrects their alpha level during multiple testing but does not make an inference about a union alternative hypothesis. In the present article, I discuss this inconsistent correction problem, including its reduction in statistical power for tests of individual hypotheses and its potential causes vis-a-vis error rate confusions and the alpha adjustment ritual. I also provide three illustrations of inconsistent corrections from recent psychology studies. I conclude that inconsistent corrections represent a symptom of statisticism, and I call for a more nuanced inference-based approach to multiple testing corrections.
翻译:在进行多重检验时,研究者常调整显著性水平以控制族系误差率,从而对联合备择假设(如"H1,1或H1,2")进行统计推断。然而在某些情况下,研究者并未进行此类联合推断,而是对构成联合假设的每个单独假设(如H1,1与H1,2)分别做出推断。例如:研究者采用Bonferroni校正将显著性水平从常规的0.050调整至0.025,分别检验H1,1与H1,2,发现H1,1显著(p<0.025)而H1,2不显著(p>0.025),由此声称支持H1,1而非H1,2。但这类单独推断并不需要调整显著性水平——唯有关于并集备择假设"H1,1或H1,2"的统计推断才需调整,因其基于两项检验中"至少一项"显著的结果,涉及族系误差率。因此,当研究者在多重检验中校正了显著性水平,却未对并集备择假设进行推断时,便产生了不一致校正。本文探讨此问题,包括其对单个假设检验统计功效的损害、误差率混淆与显著性水平校正惯例等潜在成因,并列举近期心理学研究中的三个不一致校正实例。最终指出,不一致校正是统计主义的表现,呼吁采用更精细的基于推断的多重检验校正方法。