Learning holistic computational representations in physical, chemical or biological systems requires the ability to process information from different distributions and modalities within the same model. Thus, the demand for multimodal machine learning models has sharply risen for modalities that go beyond vision and language, such as sequences, graphs, time series, or tabular data. While there are many available multimodal fusion and alignment approaches, most of them require end-to-end training, scale quadratically with the number of modalities, cannot handle cases of high modality imbalance in the training set, or are highly topology-specific, making them too restrictive for many biomedical learning tasks. This paper presents Multimodal Lego (MM-Lego), a modular and general-purpose fusion and model merging framework to turn any set of encoders into a competitive multimodal model with no or minimal fine-tuning. We achieve this by introducing a wrapper for unimodal encoders that enforces lightweight dimensionality assumptions between modalities and harmonises their representations by learning features in the frequency domain to enable model merging with little signal interference. We show that MM-Lego 1) can be used as a model merging method which achieves competitive performance with end-to-end fusion models without any fine-tuning, 2) can operate on any unimodal encoder, and 3) is a model fusion method that, with minimal fine-tuning, achieves state-of-the-art results on six benchmarked multimodal biomedical tasks.
翻译:学习物理、化学或生物系统中的整体计算表征,要求模型能够处理来自不同分布和模态的信息。因此,对于超越视觉与语言模态(如序列、图、时间序列或表格数据)的多模态机器学习模型的需求急剧增长。尽管已有多种多模态融合与对齐方法,但大多数需要端到端训练、模态数量增加时计算复杂度呈二次方增长、无法处理训练集中模态高度不平衡的情况,或高度依赖特定拓扑结构,使其在许多生物医学学习任务中限制过多。本文提出多模态乐高(MM-Lego),一种模块化通用融合与模型合并框架,可将任意编码器集合转化为具备竞争力的多模态模型,且无需或仅需极少微调。我们通过为单模态编码器引入封装层实现这一目标,该封装层强制模态间轻量级维度假设,并通过在频域学习特征以协调其表征,从而在信号干扰极小的情况下实现模型合并。我们证明MM-Lego具备以下特性:1)可作为模型合并方法,在无需任何微调的情况下达到与端到端融合模型相当的性能;2)可兼容任意单模态编码器;3)作为模型融合方法,经极少微调即可在六项基准多模态生物医学任务上取得最先进的结果。