A Bayesian approach to machine learning is attractive when we need to quantify uncertainty, deal with missing observations, when samples are scarce, or when the data is sparse. All of these commonly apply when analysing healthcare data. To address these analytical requirements, we propose a deep generative model for multinomial count data where both the weights and hidden units of the network are Dirichlet distributed. A Gibbs sampling procedure is formulated that takes advantage of a series of augmentation relations, analogous to the Zhou-Cong-Chen model. We apply the model on small handwritten digits, and a large experimental dataset of DNA mutations in cancer, and we show how the model is able to extract biologically meaningful meta-signatures in a fully data-driven way.
翻译:贝叶斯方法在机器学习中具有吸引力,尤其是在需要量化不确定性、处理缺失观测值、样本稀缺或数据稀疏的情况下。这些情况在分析医疗保健数据时普遍存在。为满足这些分析需求,我们提出了一种用于多项计数数据的深度生成模型,其中网络的权重和隐藏单元均服从狄利克雷分布。我们制定了一种吉布斯采样程序,该程序利用一系列增广关系,类似于周聪陈模型。我们将该模型应用于小规模手写数字数据集以及一个大型实验性癌症DNA突变数据集,结果表明该模型能够以完全数据驱动的方式提取具有生物学意义的元特征。