This paper is concerned with sample size determination methodology for prediction models. We propose combining the individual calculations via a learning-type curve. We suggest two distinct ways of doing so, a deterministic skeleton of a learning curve and a Gaussian process centred upon its deterministic counterpart. We employ several learning algorithms for modelling the primary endpoint and distinct measures for trial efficacy. We find that the performance may vary with the sample size, but borrowing information across sample size universally improves the performance of such calculations. The Gaussian process-based learning curve appears more robust and statistically efficient, while computational efficiency is comparable. We suggest that anchoring against historical evidence when extrapolating sample sizes should be adopted when such data are available. The methods are illustrated on binary and survival endpoints.
翻译:本文研究预测模型的样本量确定方法。我们提出通过结合学习型曲线进行个体计算。我们建议两种不同的实现方式:一种基于学习曲线的确定性骨架,另一种以确定性曲线为中心的高斯过程。我们采用多种学习算法对主要终点进行建模,并使用不同的试验效能度量指标。研究发现,模型性能可能随样本量变化而变化,但跨样本量的信息借用普遍提升了此类计算的性能。基于高斯过程的学习曲线表现出更强的鲁棒性和统计效率,而计算效率相当。我们建议在可获得历史数据时,外推样本量应基于历史证据进行锚定。本文通过二元终点和生存终点数据对方法进行了示例说明。