Deep generative models have demonstrated successful applications in learning non-linear data distributions through a number of latent variables and these models use a nonlinear function (generator) to map latent samples into the data space. On the other hand, the nonlinearity of the generator implies that the latent space shows an unsatisfactory projection of the data space, which results in poor representation learning. This weak projection, however, can be addressed by a Riemannian metric, and we show that geodesics computation and accurate interpolations between data samples on the Riemannian manifold can substantially improve the performance of deep generative models. In this paper, a Variational spatial-Transformer AutoEncoder (VTAE) is proposed to minimize geodesics on a Riemannian manifold and improve representation learning. In particular, we carefully design the variational autoencoder with an encoded spatial-Transformer to explicitly expand the latent variable model to data on a Riemannian manifold, and obtain global context modelling. Moreover, to have smooth and plausible interpolations while traversing between two different objects' latent representations, we propose a geodesic interpolation network different from the existing models that use linear interpolation with inferior performance. Experiments on benchmarks show that our proposed model can improve predictive accuracy and versatility over a range of computer vision tasks, including image interpolations, and reconstructions.
翻译:深度生成模型通过学习非线性数据分布展示了成功应用,这些模型利用潜在变量通过非线性函数(生成器)将潜在样本映射到数据空间。然而,生成器的非线性特性导致潜在空间对数据空间的投影效果不佳,从而产生较差的表示学习。但这种投影缺陷可通过黎曼度量加以解决,我们证明,在黎曼流形上进行测地线计算和数据样本间的精确插值能显著提升深度生成模型的性能。本文提出一种变分空间-Transformer自编码器(VTAE),通过最小化黎曼流形上的测地线来改进表示学习。具体而言,我们精心设计了带编码空间Transformer的变分自编码器,将潜在变量模型显式扩展到黎曼流形上的数据,并实现全局上下文建模。此外,为在两个不同对象的潜在表示间进行平滑且合理的插值,我们提出了一种测地线插值网络,不同于现有模型采用性能较差的线性插值方法。基准实验表明,我们提出的模型能在图像插值与重建等一系列计算机视觉任务中提升预测准确性和泛化能力。