Many countries have established population-based biobanks, which are being used increasingly in epidemiolgical and clinical research. These biobanks offer opportunities for large-scale studies addressing questions beyond the scope of traditional clinical trials or cohort studies. However, using biobank data poses new challenges. Typically, biobank data is collected from a study cohort recruited over a defined calendar period, with subjects entering the study at various ages falling between $c_L$ and $c_U$. This work focuses on biobank data with individuals reporting disease-onset age upon recruitment, termed prevalent data, along with individuals initially recruited as healthy, and their disease onset observed during the follow-up period. We propose a novel cumulative incidence function (CIF) estimator that efficiently incorporates prevalent cases, in contrast to existing methods, providing two advantages: (1) increased efficiency, and (2) CIF estimation for ages before the lower limit, $c_L$.
翻译:许多国家已建立基于人群的生物库,这些数据在流行病学和临床研究中得到越来越广泛的应用。生物库为开展大规模研究提供了机遇,可解决传统临床试验或队列研究范围之外的问题。然而,利用生物库数据也带来了新的挑战。通常,生物库数据来源于在特定日历期限内招募的研究队列,受试者进入研究的年龄各异,范围在c_L至c_U之间。本研究聚焦于生物库数据,其中包含招募时报告疾病发病年龄的个体(称为患病数据),以及最初作为健康个体招募并在随访期间观察到疾病发病的个体。我们提出一种新的累积发病率函数(CIF)估计方法,该方法有效整合了患病病例,与现有方法相比具有两个优势:(1)提高效率;(2)可实现下限年龄c_L之前的CIF估计。