In recent years, there has been a growing interest in the prediction of individualized treatment effects. While there is a rapidly growing literature on the development of such models, there is little literature on the evaluation of their performance. In this paper, we aim to facilitate the validation of prediction models for individualized treatment effects. The estimands of interest are defined as based on the potential outcomes framework, which facilitates a comparison of existing and novel measures. In particular, we examine existing measures of measures of discrimination for benefit (variations of the c-for-benefit), and propose model-based extensions to the treatment effect setting for discrimination and calibration metrics that have a strong basis in outcome risk prediction. The main focus is on randomized trial data with binary endpoints and on models that provide individualized treatment effect predictions and potential outcome predictions. We use simulated data to provide insight into the characteristics of the examined discrimination and calibration statistics under consideration, and further illustrate all methods in a trial of acute ischemic stroke treatment. The results show that the proposed model-based statistics had the best characteristics in terms of bias and accuracy. While resampling methods adjusted for the optimism of performance estimates in the development data, they had a high variance across replications that limited their accuracy. Therefore, individualized treatment effect models are best validated in independent data. To aid implementation, a software implementation of the proposed methods was made available in R.
翻译:近年来,个体化治疗效果预测日益受到关注。尽管此类模型的开发文献迅速增长,但关于其性能评估的研究却相对匮乏。本文旨在促进个体化治疗效果预测模型的验证。基于潜在结果框架定义感兴趣的估计量,有助于比较现有及新型评估指标。具体而言,我们考察了现有获益区分度指标(c-for-benefit的变体),并提出了基于模型的扩展方法,将结局风险预测领域具有坚实基础的区分度与校准度指标应用于治疗效果场景。研究重点聚焦于二元结局的随机试验数据,以及能提供个体化治疗效果预测和潜在结局预测的模型。我们利用模拟数据深入分析所考察的区分度与校准度统计量的特征,并在急性缺血性脑卒中治疗试验中进一步阐释所有方法。结果表明,所提出的基于模型的统计量在偏差和准确性方面具有最佳特性。尽管重采样方法能调整开发数据中性能估计的乐观性,但其在重复抽样中方差较大,限制了准确性。因此,个体化治疗效果模型最好在独立数据中进行验证。为促进实施,本文在R语言中提供了所提方法的软件实现。