Cancer, a leading cause of death globally, occurs due to genomic changes and manifests heterogeneously across patients. To advance research on personalized treatment strategies, the effectiveness of various drugs on cells derived from cancers (`cell lines') is experimentally determined in laboratory settings. Nevertheless, variations in the distribution of genomic data and drug responses between cell lines and humans arise due to biological and environmental differences. Moreover, while genomic profiles of many cancer patients are readily available, the scarcity of corresponding drug response data limits the ability to train machine learning models that can predict drug response in patients effectively. Recent cancer drug response prediction methods have largely followed the paradigm of unsupervised domain-invariant representation learning followed by a downstream drug response classification step. Introducing supervision in both stages is challenging due to heterogeneous patient response to drugs and limited drug response data. This paper addresses these challenges through a novel representation learning method in the first phase and weak supervision in the second. Experimental results on real patient data demonstrate the efficacy of our method (WISER) over state-of-the-art alternatives on predicting personalized drug response.
翻译:癌症作为全球主要死因之一,由基因组变化引发,且在不同患者间呈现异质性。为推进个性化治疗策略研究,实验室通过实验测定了多种药物对癌症来源细胞系(cell lines)的疗效。然而,由于生物学和环境差异,细胞系与人类之间的基因组数据分布及药物反应存在差异。此外,虽然大量癌症患者的基因组图谱易于获取,但相应药物反应数据的稀缺限制了训练能有效预测患者药物反应的机器学习模型的能力。近期癌症药物反应预测方法主要遵循先进行无监督域不变表示学习、后进行下游药物反应分类的范式。由于患者对药物的异质性反应以及药物反应数据有限,在两个阶段引入监督学习均面临挑战。本文通过第一阶段的新型表示学习方法和第二阶段的弱监督学习应对这些挑战。基于真实患者数据的实验结果表明,我们提出的方法(WISER)在预测个性化药物反应方面优于当前最先进的替代方法。