In the absence of artificial labels, the independent and dependent features in the data are cluttered. How to construct the inductive biases of the model to flexibly divide and effectively contain features with different complexity is the main focal point of unsupervised disentangled representation learning. This paper proposes a new iterative decomposition path of total correlation and explains the disentangled representation ability of VAE from the perspective of model capacity allocation. The newly developed objective function combines latent variable dimensions into joint distribution while relieving the independence constraints of marginal distributions in combination, leading to latent variables with a more manipulable prior distribution. The novel model enables VAE to adjust the parameter capacity to divide dependent and independent data features flexibly. Experimental results on various datasets show an interesting relevance between model capacity and the latent variable grouping size, called the "V"-shaped best ELBO trajectory. Additionally, we empirically demonstrate that the proposed method obtains better disentangling performance with reasonable parameter capacity allocation.
翻译:在缺乏人工标签的情况下,数据中的独立特征与依赖特征相互混杂。如何构建模型的归纳偏好,以灵活划分并有效容纳不同复杂度的特征,是无监督解耦表示学习的核心关注点。本文提出了一种新的总相关性迭代分解路径,并从模型容量分配的角度解释了VAE的解耦表示能力。新开发的目标函数将潜变量维度组合为联合分布,同时放宽组合中边际分布的独立性约束,从而获得具有更可控先验分布的潜变量。该新型模型使VAE能够调整参数容量,灵活划分数据的依赖特征与独立特征。在多种数据集上的实验结果表明,模型容量与潜变量分组大小之间存在一种有趣的关联,称为"V"形最佳ELBO轨迹。此外,我们通过实验证明,所提方法通过合理的参数容量分配能够获得更优的解耦性能。