Multimodal contrastive representation learning methods have proven successful across a range of domains, partly due to their ability to generate meaningful shared representations of complex phenomena. To enhance the depth of analysis and understanding of these acquired representations, we introduce a unified causal model specifically designed for multimodal data. By examining this model, we show that multimodal contrastive representation learning excels at identifying latent coupled variables within the proposed unified model, up to linear or permutation transformations resulting from different assumptions. Our findings illuminate the potential of pre-trained multimodal models, eg, CLIP, in learning disentangled representations through a surprisingly simple yet highly effective tool: linear independent component analysis. Experiments demonstrate the robustness of our findings, even when the assumptions are violated, and validate the effectiveness of the proposed method in learning disentangled representations.
翻译:多模态对比表示学习方法已在多个领域取得成功,部分原因在于其能够生成复杂现象的有意义共享表示。为深化对这些习得表示的分析与理解,我们引入了一个专门针对多模态数据的统一因果模型。通过分析该模型,我们表明多模态对比表示学习能够有效识别所提出统一模型中的潜在耦合变量,其精度可达因不同假设而产生的线性变换或排列变换。我们的发现揭示了预训练多模态模型(如CLIP)在通过一个极其简单却高效的工具——线性独立成分分析——学习解耦表示方面的潜力。实验证明,即使假设条件被违背,我们的结论仍具有鲁棒性,并验证了所提方法在学习解耦表示方面的有效性。