The intersection of chemistry and Artificial Intelligence (AI) is an active area of research focused on accelerating scientific discovery. While using large language models (LLMs) with scientific modalities has shown potential, there are significant challenges to address, such as improving training efficiency and dealing with the out-of-distribution problem. Focussing on the task of automated language-molecule translation, we are the first to use state-of-the art (SOTA) human-centric optimisation algorithms in the cross-modal setting, successfully aligning cross-language-molecule modals. We empirically show that we can augment the capabilities of scientific LLMs without the need for extensive data or large models. We conduct experiments using only 10% of the available data to mitigate memorisation effects associated with training large models on extensive datasets. We achieve significant performance gains, surpassing the best benchmark model trained on extensive in-distribution data by a large margin and reach new SOTA levels. Additionally we are the first to propose employing non-linear fusion for mixing cross-modal LLMs which further boosts performance gains without increasing training costs or data needs. Finally, we introduce a fine-grained, domain-agnostic evaluation method to assess hallucination in LLMs and promote responsible use.
翻译:化学与人工智能(AI)的交叉领域是一个活跃的研究方向,旨在加速科学发现。虽然将大语言模型(LLMs)应用于科学模态已显示出潜力,但仍存在一些重大挑战需要解决,例如提高训练效率和处理分布外问题。聚焦于自动化语言-分子翻译任务,我们首次在跨模态设置中使用了最先进的以人为中心的优化算法,成功对齐了跨语言-分子模态。我们通过实验证明,无需大量数据或大型模型即可增强科学大语言模型的能力。我们仅使用10%的可用数据进行实验,以减轻在广泛数据集上训练大型模型所带来的记忆效应。我们取得了显著的性能提升,大幅超越了在大量分布内数据上训练的最佳基准模型,并达到了新的最先进水平。此外,我们首次提出采用非线性融合方法来混合跨模态大语言模型,这在不增加训练成本或数据需求的情况下进一步提升了性能增益。最后,我们引入了一种细粒度的、领域无关的评估方法,用于评估大语言模型中的幻觉问题,并促进其负责任的使用。