In statistics, it is important to have realistic data sets available for a particular context to allow an appropriate and objective method comparison. For many use cases, benchmark data sets for method comparison are already available online. However, in most medical applications and especially for clinical trials in oncology, there is a lack of adequate benchmark data sets, as patient data can be sensitive and therefore cannot be published. A potential solution for this are simulation studies. However, it is sometimes not clear, which simulation models are suitable for generating realistic data. A challenge is that potentially unrealistic assumptions have to be made about the distributions. Our approach is to use reconstructed benchmark data sets %can be used as a basis for the simulations, which has the following advantages: the actual properties are known and more realistic data can be simulated. There are several possibilities to simulate realistic data from benchmark data sets. We investigate simulation models based upon kernel density estimation, fitted distributions, case resampling and conditional bootstrapping. In order to make recommendations on which models are best suited for a specific survival setting, we conducted a comparative simulation study. Since it is not possible to provide recommendations for all possible survival settings in a single paper, we focus on providing realistic simulation models for two-armed phase III lung cancer studies. To this end we reconstructed benchmark data sets from recent studies. We used the runtime and different accuracy measures (effect sizes and p-values) as criteria for comparison.
翻译:在统计学中,针对特定情境拥有真实数据集对于进行适当且客观的方法比较至关重要。许多用例的方法比较基准数据集已可在线获取。然而,在大多数医学应用中,尤其是肿瘤学临床试验中,由于患者数据具有敏感性而无法公开发布,因此缺乏充分的基准数据集。为此,仿真研究成为一种潜在解决方案。但有时尚不明确哪些模拟模型适用于生成真实数据。其挑战在于,往往需要对数据分布做出可能不切实际的假设。我们的方法是以重建的基准数据集作为模拟基础,具有以下优势:实际特征已知,且可模拟出更真实的数据。从基准数据集模拟真实数据存在多种可能性。我们研究了基于核密度估计、拟合分布、案例重抽样及条件自助法的模拟模型。为确定在特定生存情境中最优的模型选择,我们开展了一项比较仿真研究。由于无法在一篇论文中为所有可能的生存情境提供建议,我们聚焦于为双臂III期肺癌研究提供真实模拟模型。为此,我们重建了近期研究的基准数据集,并以运行时间和不同精度指标(效应量与p值)作为比较标准。