A Sparse Graph-Structured Lasso Mixed Model for Genetic Association with Confounding Correction

While linear mixed model (LMM) has shown a competitive performance in correcting spurious associations raised by population stratification, family structures, and cryptic relatedness, more challenges are still to be addressed regarding the complex structure of genotypic and phenotypic data. For example, geneticists have discovered that some clusters of phenotypes are more co-expressed than others. Hence, a joint analysis that can utilize such relatedness information in a heterogeneous data set is crucial for genetic modeling. We proposed the sparse graph-structured linear mixed model (sGLMM) that can incorporate the relatedness information from traits in a dataset with confounding correction. Our method is capable of uncovering the genetic associations of a large number of phenotypes together while considering the relatedness of these phenotypes. Through extensive simulation experiments, we show that the proposed model outperforms other existing approaches and can model correlation from both population structure and shared signals. Further, we validate the effectiveness of sGLMM in the real-world genomic dataset on two different species from plants and humans. In Arabidopsis thaliana data, sGLMM behaves better than all other baseline models for 63.4% traits. We also discuss the potential causal genetic variation of Human Alzheimer's disease discovered by our model and justify some of the most important genetic loci.

翻译：虽然线性混合模型在纠正由群体分层、家族结构和隐秘亲缘关系引起的虚假关联方面表现出色，但针对基因型和表型数据的复杂结构仍存在诸多挑战。例如，遗传学家发现某些表型簇的共表达程度高于其他簇。因此，在异质性数据集中利用这种关联信息进行联合分析对遗传建模至关重要。我们提出了稀疏图结构线性混合模型，该模型能够在混杂校正下整合数据集中性状间的关联信息。该方法能够同时揭示大量表型的遗传关联，同时考虑这些表型的相关性。通过广泛的模拟实验，我们证明所提出的模型优于其他现有方法，能够同时建模群体结构和共享信号引起的相关性。此外，我们利用来自植物和人类两个不同物种的真实基因组数据集验证了sGLMM的有效性。在拟南芥数据中，sGLMM在63.4%的性状上优于所有其他基线模型。我们还讨论了模型发现的人类阿尔茨海默病潜在因果遗传变异，并验证了其中一些最重要的遗传位点。

相关内容

MoDELS

关注 45

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

不可错过！杜克大学《因果推断》课程，全面讲述因果推理

专知会员服务

52+阅读 · 2022年10月22日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

【ETH】最新《几何数据分析》2020课程，附PPT下载

专知会员服务

45+阅读 · 2020年12月18日

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

19+阅读 · 2019年10月22日