Recent genome-wide association studies (GWAS) have uncovered the genetic basis of complex traits, but show an under-representation of non-European descent individuals, underscoring a critical gap in genetic research. Here, we assess whether we can improve disease prediction across diverse ancestries using multiomic data. We evaluate the performance of Group-LASSO INTERaction-NET (glinternet) and pretrained lasso in disease prediction focusing on diverse ancestries in the UK Biobank. Models were trained on data from White British and other ancestries and validated across a cohort of over 96,000 individuals for 8 diseases. Out of 96 models trained, we report 16 with statistically significant incremental predictive performance in terms of ROC-AUC scores (p-value < 0.05), found for diabetes, arthritis, gall stones, cystitis, asthma and osteoarthritis. For the interaction and pretrained models that outperformed the baseline, the PRS score was the primary driver behind prediction. Our findings indicate that both interaction terms and pre-training can enhance prediction accuracy but for a limited set of diseases and moderate improvements in accuracy
翻译:近期全基因组关联研究已揭示复杂性状的遗传基础,但非欧洲血统个体代表性不足,凸显了遗传研究的关键空白。本研究评估了利用多组学数据改善跨祖源疾病预测的可行性。我们评估了Group-LASSO INTERaction-NET(glinternet)和预训练lasso在UK Biobank中针对不同祖源人群的疾病预测性能。模型基于英国白人及其他祖源数据进行训练,并在超过96,000名个体的队列中对8种疾病进行验证。在训练的96个模型中,我们报告了16个模型在ROC-AUC评分方面具有统计显著的增量预测性能(p值<0.05),涵盖糖尿病、关节炎、胆结石、膀胱炎、哮喘和骨关节炎。对于优于基线的交互模型和预训练模型,PRS评分是预测的主要驱动因素。研究结果表明,交互项和预训练均可提升预测精度,但仅适用于有限疾病且精度提升幅度适中。