Understanding the biological mechanisms of disease is crucial for medicine, and in particular, for drug discovery. AI-powered analysis of genome-scale biological data holds great potential in this regard. The increasing availability of single-cell RNA sequencing data has enabled the development of large foundation models for disease biology. However, existing foundation models only modestly improve over task-specific models in downstream applications. Here, we explored two avenues for improving single-cell foundation models. First, we scaled the pre-training data to a diverse collection of 116 million cells, which is larger than those used by previous models. Second, we leveraged the availability of large-scale biological annotations as a form of supervision during pre-training. We trained the \model family of models comprising six transformer-based state-of-the-art single-cell foundation models with 70 million, 160 million, and 400 million parameters. We vetted our models on several downstream evaluation tasks, including identifying the underlying disease state of held-out donors not seen during training, distinguishing between diseased and healthy cells for disease conditions and donors not seen during training, and probing the learned representations for known biology. Our models showed substantial improvement over existing works, and scaling experiments showed that performance improved predictably with both data volume and parameter count.
翻译:理解疾病的生物学机制对医学至关重要,尤其在新药研发领域。基于人工智能的全基因组规模生物学数据分析在此方面具有巨大潜力。单细胞RNA测序数据日益丰富的可及性,推动了面向疾病生物学的大型基础模型开发。然而,现有基础模型在下游任务中的性能提升幅度有限,仅略优于任务特定模型。为此,我们探索了两条改进单细胞基础模型的路径:首先,将预训练数据规模扩展至包含1.16亿个细胞的多样化集合,该规模超越以往所有模型;其次,利用大规模生物学注释作为预训练阶段的监督信号。我们训练了包含六个基于Transformer架构的当前最先进单细胞基础模型的模型系列,参数量分别为7000万、1.6亿和4亿。通过在多个下游评估任务中的验证,包括识别训练中未见过的保留供体的疾病状态、区分训练未涉及疾病条件及供体的病变与健康细胞,以及探测模型表征中蕴含的已知生物学知识,我们的模型展现出相较于现有工作的显著提升。规模化实验表明,模型性能随数据量和参数量的增加呈可预测性提升。