The posterior collapse phenomenon in variational autoencoder (VAE), where the variational posterior distribution closely matches the prior distribution, can hinder the quality of the learned latent variables. As a consequence of posterior collapse, the latent variables extracted by the encoder in VAE preserve less information from the input data and thus fail to produce meaningful representations as input to the reconstruction process in the decoder. While this phenomenon has been an actively addressed topic related to VAE performance, the theory for posterior collapse remains underdeveloped, especially beyond the standard VAE. In this work, we advance the theoretical understanding of posterior collapse to two important and prevalent yet less studied classes of VAE: conditional VAE and hierarchical VAE. Specifically, via a non-trivial theoretical analysis of linear conditional VAE and hierarchical VAE with two levels of latent, we prove that the cause of posterior collapses in these models includes the correlation between the input and output of the conditional VAE and the effect of learnable encoder variance in the hierarchical VAE. We empirically validate our theoretical findings for linear conditional and hierarchical VAE and demonstrate that these results are also predictive for non-linear cases with extensive experiments.
翻译:变分自编码器(VAE)中的后验崩塌现象——即变分后验分布与先验分布高度吻合——会损害学习到的潜在变量的质量。由于后验崩塌,编码器提取的潜在变量保留的输入数据信息减少,从而无法生成有意义的表征用于解码器的重构过程。尽管这一现象已成为影响VAE性能的关键研究课题,但其理论机制仍不完善,尤其针对标准VAE之外的模型。本研究将后验崩塌的理论理解推进至两类重要且普遍、但研究较少的VAE变体:条件VAE与分层VAE。具体而言,通过对线性条件VAE与包含两层潜在变量的线性分层VAE进行非平凡的理论分析,我们证明这两类模型中后验崩塌的成因包括:条件VAE中输入与输出之间的相关性,以及分层VAE中可学习编码器方差的影响。我们通过实验验证了线性条件VAE与分层VAE的理论发现,并通过大量实验证明这些结论对非线性情形同样具有预测性。