The remarkable success of Large Language Models (LLMs) across diverse tasks has driven the research community to extend their capabilities to molecular applications. However, most molecular LLMs employ adapter-based architectures that do not treat molecule and text modalities equally and lack a supervision signal for the molecule modality. To address these issues, we introduce UniMoT, a Unified Molecule-Text LLM adopting a tokenizer-based architecture that expands the vocabulary of LLM with molecule tokens. Specifically, we introduce a Vector Quantization-driven tokenizer that incorporates a Q-Former to bridge the modality gap between molecule and text. This tokenizer transforms molecules into sequences of molecule tokens with causal dependency, encapsulating high-level molecular and textual information. Equipped with this tokenizer, UniMoT can unify molecule and text modalities under a shared token representation and an autoregressive training paradigm, enabling it to interpret molecules as a foreign language and generate them as text. Following a four-stage training scheme, UniMoT emerges as a multi-modal generalist capable of performing both molecule-to-text and text-to-molecule tasks. Extensive experiments demonstrate that UniMoT achieves state-of-the-art performance across a wide range of molecule comprehension and generation tasks.
翻译:大型语言模型(LLM)在多样化任务中取得的显著成功,推动了研究界将其能力扩展至分子应用领域。然而,大多数分子LLM采用基于适配器的架构,未能平等对待分子与文本模态,且缺乏针对分子模态的监督信号。为解决这些问题,我们提出了UniMoT——一种采用基于分词器架构的统一分子-文本LLM,通过引入分子令牌来扩展LLM的词表。具体而言,我们设计了一种基于向量量化的分词器,该分词器整合了Q-Former以弥合分子与文本之间的模态鸿沟。该分词器将分子转化为具有因果依赖关系的分子令牌序列,从而封装了高层级的分子与文本信息。借助此分词器,UniMoT能够在共享的令牌表示和自回归训练范式下统一分子与文本模态,使其能够将分子解析为一种"外语"并以文本形式生成分子。通过四阶段训练方案,UniMoT最终成为能够同时执行分子到文本与文本到分子任务的多模态通用模型。大量实验表明,UniMoT在广泛的分子理解与生成任务中均达到了最先进的性能水平。