Advances in single-cell omics allow for unprecedented insights into the transcription profiles of individual cells. When combined with large-scale perturbation screens, through which specific biological mechanisms can be targeted, these technologies allow for measuring the effect of targeted perturbations on the whole transcriptome. These advances provide an opportunity to better understand the causative role of genes in complex biological processes such as gene regulation, disease progression or cellular development. However, the high-dimensional nature of the data, coupled with the intricate complexity of biological systems renders this task nontrivial. Within the machine learning community, there has been a recent increase of interest in causality, with a focus on adapting established causal techniques and algorithms to handle high-dimensional data. In this perspective, we delineate the application of these methodologies within the realm of single-cell genomics and their challenges. We first present the model that underlies most of current causal approaches to single-cell biology and discuss and challenge the assumptions it entails from the biological point of view. We then identify open problems in the application of causal approaches to single-cell data: generalising to unseen environments, learning interpretable models, and learning causal models of dynamics. For each problem, we discuss how various research directions - including the development of computational approaches and the adaptation of experimental protocols - may offer ways forward, or on the contrary pose some difficulties. With the advent of single cell atlases and increasing perturbation data, we expect causal models to become a crucial tool for informed experimental design.
翻译:单细胞组学技术的进步使我们能够以前所未有的分辨率洞察单个细胞的转录图谱。当与大规模扰动筛选(可靶向特定生物学机制)相结合时,这些技术能够测量靶向扰动对整个转录组的影响。这些进展为深入理解基因在基因调控、疾病进展或细胞发育等复杂生物学过程中的因果作用提供了契机。然而,数据的高维特性与生物系统的复杂交织性使这一任务颇具挑战。近年来,机器学习领域对因果性的关注度显著提升,重点在于改进既有因果技术与算法以处理高维数据。本文从研究视角出发,系统阐述这些方法在单细胞基因组学领域的应用及其面临的挑战。我们首先介绍当前大多数单细胞生物学因果方法所依赖的基础模型,并从生物学角度讨论并质疑其隐含假设。随后,我们识别出因果方法应用于单细胞数据时的若干开放性问题:泛化至未知环境、学习可解释模型以及学习动力学因果模型。针对每个问题,我们探讨不同研究方向(包括计算方法开发与实验方案改进)如何提供突破契机或可能带来新困难。随着单细胞图谱的涌现与扰动数据的不断积累,我们预期因果模型将成为指导实验设计的关键工具。