Multimodal representation learning techniques typically rely on paired samples to learn common representations, but paired samples are challenging to collect in fields such as biology where measurement devices often destroy the samples. This paper presents an approach to address the challenge of aligning unpaired samples across disparate modalities in multimodal representation learning. We draw an analogy between potential outcomes in causal inference and potential views in multimodal observations, which allows us to use Rubin's framework to estimate a common space in which to match samples. Our approach assumes we collect samples that are experimentally perturbed by treatments, and uses this to estimate a propensity score from each modality, which encapsulates all shared information between a latent state and treatment and can be used to define a distance between samples. We experiment with two alignment techniques that leverage this distance -- shared nearest neighbours (SNN) and optimal transport (OT) matching -- and find that OT matching results in significant improvements over state-of-the-art alignment approaches in both a synthetic multi-modal setting and in real-world data from NeurIPS Multimodal Single-Cell Integration Challenge.
翻译:多模态表征学习技术通常依赖于配对样本来学习共同表征,但在生物学等领域,由于测量设备常会破坏样本,因此收集配对样本具有挑战性。本文提出了一种方法,以解决多模态表征学习中跨不同模态对齐非配对样本的难题。我们将因果推断中的潜在结果与多模态观测中的潜在视图进行类比,从而利用鲁宾框架估计一个用于样本匹配的共同空间。我们的方法假设收集了经实验性处理扰动的样本,并据此从每种模态中估计出一个倾向性分数。该分数封装了潜在状态与处理之间的所有共享信息,可用于定义样本间的距离。我们实验了两种利用这一距离的对齐技术——共享最近邻和最优传输匹配——并发现最优传输匹配在合成多模态设置及NeurIPS多模态单细胞整合挑战赛的真实世界数据中,相较于最先进的对齐方法取得了显著改进。