Variational Autoencoder (VAE)-based generative models offer flexible representation learning by incorporating meta-priors, general premises considered beneficial for downstream tasks. However, the incorporated meta-priors often involve ad-hoc model deviations from the original likelihood architecture, causing undesirable changes in their training. In this paper, we propose a novel representation learning method, Gromov-Wasserstein Autoencoders (GWAE), which directly matches the latent and data distributions using the variational autoencoding scheme. Instead of likelihood-based objectives, GWAE models minimize the Gromov-Wasserstein (GW) metric between the trainable prior and given data distributions. The GW metric measures the distance structure-oriented discrepancy between distributions even with different dimensionalities, which provides a direct measure between the latent and data spaces. By restricting the prior family, we can introduce meta-priors into the latent space without changing their objective. The empirical comparisons with VAE-based models show that GWAE models work in two prominent meta-priors, disentanglement and clustering, with their GW objective unchanged.
翻译:基于变分自编码器(VAE)的生成模型通过引入元先验(即被认为有益于下游任务的通用前提)提供了灵活的表征学习。然而,这些元先验往往会导致模型在原始似然架构上产生临时性偏离,从而对其训练过程造成非期望的改变。本文提出了一种新型表征学习方法——Gromov-Wasserstein自编码器(GWAE),该方法利用变分自编码方案直接匹配隐空间分布与数据分布。不同于基于似然的目标函数,GWAE模型通过最小化可训练先验分布与给定数据分布之间的Gromov-Wasserstein(GW)度量来运作。该GW度量能够衡量不同维度分布之间基于距离结构导向的差异,从而为隐空间与数据空间提供直接度量。通过限制先验分布族,我们可在不改变目标函数的前提下将元先验引入隐空间。与基于VAE模型的实验比较表明,GWAE模型在解耦与聚类这两个重要元先验任务中,其GW目标函数始终保持不变。