Precision cancer medicine aims to determine the optimal treatment for each patient. In-vitro cancer drug sensitivity screens combined with multi-omics characterization of the cancer cells have become an important tool to achieve this aim. Analyzing such pharmacogenomic studies requires flexible and efficient joint statistical models for associating drug sensitivity with high-dimensional multi-omics data. We propose a multivariate Bayesian structured variable selection model for sparse identification of omics features associated with multiple correlated drug responses. Since many anti-cancer drugs are designed for specific molecular targets, our approach makes use of known structure between responses and predictors, e.g. molecular pathways and related omics features targeted by specific drugs, via a Markov random field (MRF) prior for the latent indicator variables of the coefficients in sparse seemingly unrelated regression. The structure information included in the MRF prior can improve the model performance, i.e. variable selection and response prediction, compared to other common priors. In addition, we employ random effects to capture heterogeneity between cancer types in a pan-cancer setting. The proposed approach is validated by simulation studies and applied to the Genomics of Drug Sensitivity in Cancer data, which includes pharmacological profiling and multi-omics characterization of a large set of heterogeneous cell lines.
翻译:精准癌症医学旨在为每位患者确定最优治疗方案。结合癌细胞多组学表征的体外抗癌药物敏感性筛选已成为实现这一目标的重要工具。分析这样的药物基因组研究需要灵活且高效的联合统计模型,以将药物敏感性与高维多组学数据关联起来。我们提出了一种多元贝叶斯结构化变量选择模型,用于稀疏识别与多个相关药物反应相关联的组学特征。由于许多抗癌药物针对特定的分子靶点设计,我们的方法通过马尔可夫随机场(MRF)先验,对稀疏看似不相关回归中的系数潜在指示变量进行建模,从而利用了反应与预测变量之间的已知结构(例如,特定药物靶向的分子通路及相关组学特征)。与其他常见先验相比,MRF先验中包含的结构信息可以提高模型性能,即变量选择和反应预测。此外,我们采用随机效应来捕捉泛癌背景下不同癌症类型之间的异质性。所提出方法通过模拟研究进行了验证,并应用于《癌症药物敏感性基因组学》数据集,该数据集包含大量异质性细胞系的药理学特征和多组学表征。