Learning low-dimensional representations of single-cell transcriptomics has become instrumental to its downstream analysis. The state of the art is currently represented by neural network models such as variational autoencoders (VAEs) which use a variational approximation of the likelihood for inference. We here present the Deep Generative Decoder (DGD), a simple generative model that computes model parameters and representations directly via maximum a posteriori (MAP) estimation. The DGD handles complex parameterized latent distributions naturally unlike VAEs which typically use a fixed Gaussian distribution, because of the complexity of adding other types. We first show its general functionality on a commonly used benchmark set, Fashion-MNIST. Secondly, we apply the model to multiple single-cell data sets. Here the DGD learns low-dimensional, meaningful and well-structured latent representations with sub-clustering beyond the provided labels. The advantages of this approach are its simplicity and its capability to provide representations of much smaller dimensionality than a comparable VAE.
翻译:学习单细胞转录组学的低维表征已成为其下游分析的重要工具。当前最先进的方法由神经网络模型(如变分自编码器VAEs)代表,这类模型利用似然的变分近似进行推理。本文提出深度生成解码器(DGD),一种直接通过最大后验(MAP)估计计算模型参数和表征的简单生成模型。与通常使用固定高斯分布的VAEs不同(因其难以引入其他分布类型),DGD能自然地处理参数化潜在分布。我们首先在通用基准数据集Fashion-MNIST上展示其基础功能,进而将该模型应用于多个单细胞数据集。实验表明,DGD能够学习到低维、有意义的且结构清晰的潜在表征,并在已提供标签之外实现亚聚类。该方法的优势在于其简洁性,以及能够提供比同类VAE更低维度的表征能力。