Prediction performance metrics such as accuracy and the F1 score are typically reported as single numbers, with no measure of uncertainty. The omission has been tolerable in exploratory settings, where model evaluation is used for informal comparison rather than formal decision-making. But as machine learning is deployed in real-world applications, evaluation results are increasingly used to support binary decisions -- whether a model meets a required standard or not -- making uncertainty quantification essential. The problem is compounded when data are dependent, as in repeated measurements, clustered subjects, or time series, where variability is harder to assess and easy to underestimate. We develop a unified framework that links a broad class of performance metrics through their representation as smooth functionals of confusion-matrix probabilities. This representation allows the use of the cluster-robust sandwich variance estimator to obtain asymptotically valid confidence intervals, hypothesis tests, and paired model comparisons for both binary and multiclass problems under clustered data. We also provide power and sample size approximations based on pilot data, enabling principled study design for model evaluation. Simulations show that the proposed methods achieve near-nominal coverage across a range of dependence structures, while naive methods underestimate variability. A real-data application further illustrates how accounting for clustering can materially change conclusions. These results offer a practical foundation for uncertainty quantification and study design in prediction performance evaluation, in settings where decisions should be justified under dependent and clustered data.
翻译:预测性能度量(如准确率和F1分数)通常以单一数值报告,而不包含不确定性测度。这种省略在探索性场景中尚可容忍——其中模型评估用于非正式比较而非正式决策。然而,随着机器学习在现实应用中部署,评估结果日益被用于支持二元决策(模型是否达到所需标准),使得不确定性量化变得至关重要。当数据存在依赖性时(如重复测量、聚集受试者或时间序列数据),问题更加复杂:变异性更难评估且容易被低估。我们开发了一个统一框架,通过将广泛类别的性能度量表示为混淆矩阵概率的平滑泛函,揭示其内在联系。这种表示允许使用聚类稳健的夹心方差估计量,在聚类数据下为二元和多分类问题获得渐近有效的置信区间、假设检验和配对模型比较。我们还基于试点数据提供了统计功效和样本量近似方法,支持模型评估的原则性研究设计。模拟实验表明,所提方法在多种依赖性结构下实现接近名义覆盖率的性能,而朴素方法会低估变异性。实际数据应用进一步说明考虑聚类效应可能实质性地改变结论。这些结果为依赖性和聚类数据场景下预测性能评估中的不确定性量化和研究设计提供了实用基础。