When factorized approximations are used for variational inference (VI), they tend to underestimate the uncertainty -- as measured in various ways -- of the distributions they are meant to approximate. We consider two popular ways to measure the uncertainty deficit of VI: (i) the degree to which it underestimates the componentwise variance, and (ii) the degree to which it underestimates the entropy. To better understand these effects, and the relationship between them, we examine an informative setting where they can be explicitly (and elegantly) analyzed: the approximation of a Gaussian,~$p$, with a dense covariance matrix, by a Gaussian,~$q$, with a diagonal covariance matrix. We prove that $q$ always underestimates both the componentwise variance and the entropy of $p$, \textit{though not necessarily to the same degree}. Moreover we demonstrate that the entropy of $q$ is determined by the trade-off of two competing forces: it is decreased by the shrinkage of its componentwise variances (our first measure of uncertainty) but it is increased by the factorized approximation which delinks the nodes in the graphical model of $p$. We study various manifestations of this trade-off, notably one where, as the dimension of the problem grows, the per-component entropy gap between $p$ and $q$ becomes vanishingly small even though $q$ underestimates every componentwise variance by a constant multiplicative factor. We also use the shrinkage-delinkage trade-off to bound the entropy gap in terms of the problem dimension and the condition number of the correlation matrix of $p$. Finally we present empirical results on both Gaussian and non-Gaussian targets, the former to validate our analysis and the latter to explore its limitations.
翻译:当因子化近似用于变分推断(VI)时,它们往往会低估其试图近似的分布的不确定性(以多种方式衡量)。我们考虑了衡量VI不确定性不足的两种常用方法:(i)低估分量方差的程度,以及(ii)低估熵的程度。为了更好地理解这些效应及其相互关系,我们考察了一个信息丰富的场景,其中可以明确(且优雅地)分析这些效应:用具有对角协方差矩阵的高斯分布q来近似具有密集协方差矩阵的高斯分布p。我们证明q总是低估p的分量方差和熵,但未必达到相同的程度。此外,我们证明q的熵由两种竞争力量的权衡决定:其分量方差的收缩(我们不确定性的第一个度量)降低熵,但因子化近似(解耦p图模型中的节点)增加熵。我们研究了这种权衡的各种表现形式,特别关注当问题维度增加时,即使q以恒定的乘法因子低估每个分量方差,p与q之间的每分量熵差也会变得极小的情形。我们还利用收缩-解耦权衡来根据问题维度和p的相关系数矩阵的条件数限制熵差。最后,我们给出了高斯和非高斯目标上的实证结果,前者用于验证我们的分析,后者用于探索其局限性。