There are two distinct definitions of 'P-value' for evaluating a proposed hypothesis or model for the process generating an observed dataset. The original definition starts with a measure of the divergence of the dataset from what was expected under the model, such as a sum of squares or a deviance statistic. A P-value is then the ordinal location of the measure in a reference distribution computed from the model and the data, and is treated as a unit-scaled index of compatibility between the data and the model. In the other definition, a P-value is a random variable on the unit interval whose realizations can be compared to a cutoff alpha to generate a decision rule with known error rates under the model and specific alternatives. It is commonly assumed that realizations of such decision P-values always correspond to divergence P-values. But this need not be so: Decision P-values can violate intuitive single-sample coherence criteria where divergence P-values do not. It is thus argued that divergence and decision P-values should be carefully distinguished in teaching, and that divergence P-values are the relevant choice when the analysis goal is to summarize evidence rather than implement a decision rule.
翻译:关于评估解释观测数据集生成过程的假设或模型,存在两种不同的“P值”定义。原始定义始于衡量数据与模型预期之间偏离程度的指标,例如平方和或偏差统计量。P值是该指标在由模型和数据计算的参考分布中的序数位置,被视为数据与模型之间兼容性的单位尺度化指数。在另一种定义中,P值是单位区间上的随机变量,其实测值可与临界值alpha进行比较,以生成在模型及特定备择假设下具有已知错误率的决策规则。通常假设此类决策P值的实现值始终对应散度P值,但事实未必如此:决策P值可能违反直观的单样本一致性准则,而散度P值则不会违反。因此本文主张:在教学过程中应谨慎区分散度P值与决策P值,当分析目标是总结证据而非实施决策规则时,散度P值是相关的选择。