Spatial transcriptomics (ST) captures gene expression within distinct regions (i.e., windows) of a tissue slide. Traditional supervised learning frameworks applied to model ST are constrained to predicting expression from slide image windows for gene types seen during training, failing to generalize to unseen gene types. To overcome this limitation, we propose a semantic guided network (SGN), a pioneering zero-shot framework for predicting gene expression from slide image windows. Considering a gene type can be described by functionality and phenotype, we dynamically embed a gene type to a vector per its functionality and phenotype, and employ this vector to project slide image windows to gene expression in feature space, unleashing zero-shot expression prediction for unseen gene types. The gene type functionality and phenotype are queried with a carefully designed prompt from a pre-trained large language model (LLM). On standard benchmark datasets, we demonstrate competitive zero-shot performance compared to past state-of-the-art supervised learning approaches.
翻译:空间转录组学(ST)能够捕获组织切片中不同区域(即窗口)的基因表达。传统的监督学习框架在应用于ST建模时,只能预测训练过程中所见基因类型在切片图像窗口中的表达,无法泛化到未见过的基因类型。为突破这一限制,我们提出语义引导网络(SGN),这是一种开创性的零样本框架,可从切片图像窗口预测基因表达。考虑到基因类型可通过功能与表型进行描述,我们根据基因类型的功能和表型动态将其嵌入为向量,并利用该向量将切片图像窗口投影到特征空间中的基因表达,从而实现对未见基因类型的零样本表达预测。基因类型的功能与表型通过精心设计的提示从预训练的大语言模型(LLM)中查询得到。在标准基准数据集上,我们证明了该方法相比以往最先进的监督学习方法,能够实现具有竞争力的零样本性能。