The posterior collapse phenomenon in variational autoencoder (VAE), where the variational posterior distribution closely matches the prior distribution, can hinder the quality of the learned latent variables. As a consequence of posterior collapse, the latent variables extracted by the encoder in VAE preserve less information from the input data and thus fail to produce meaningful representations as input to the reconstruction process in the decoder. While this phenomenon has been an actively addressed topic related to VAE performance, the theory for posterior collapse remains underdeveloped, especially beyond the standard VAE. In this work, we advance the theoretical understanding of posterior collapse to two important and prevalent yet less studied classes of VAE: conditional VAE and hierarchical VAE. Specifically, via a non-trivial theoretical analysis of linear conditional VAE and hierarchical VAE with two levels of latent, we prove that the cause of posterior collapses in these models includes the correlation between the input and output of the conditional VAE and the effect of learnable encoder variance in the hierarchical VAE. We empirically validate our theoretical findings for linear conditional and hierarchical VAE and demonstrate that these results are also predictive for non-linear cases with extensive experiments.
翻译:变分自编码器中的后验坍缩现象(即变分后验分布与先验分布高度趋同)会削弱学习到的潜在变量的质量。由于后验坍缩,编码器提取的潜在变量对输入数据的信息保留能力下降,从而无法为解码器的重构过程提供有意义的表征。尽管该现象已成为变分自编码器性能研究中被广泛关注的问题,但关于后验坍缩的理论仍不完善,尤其是对于标准变分自编码器之外的模型。本研究将后验坍缩的理论理解拓展至两类重要且普遍但研究较少的变分自编码器变体:条件变分自编码器与层级变分自编码器。具体而言,通过针对线性条件变分自编码器及双层潜在变量的层级变分自编码器开展深入理论分析,我们证明这两类模型中后验坍缩的成因包括:条件变分自编码器的输入与输出之间的相关性,以及层级变分自编码器中可学习编码器方差的影响。我们通过线性条件与层级变分自编码器实验验证了理论发现,并进一步通过大量实验证明这些结论对非线性情形同样具有预测能力。