Ocean science, which delves into the oceans that are reservoirs of life and biodiversity, is of great significance given that oceans cover over 70% of our planet's surface. Recently, advances in Large Language Models (LLMs) have transformed the paradigm in science. Despite the success in other domains, current LLMs often fall short in catering to the needs of domain experts like oceanographers, and the potential of LLMs for ocean science is under-explored. The intrinsic reason may be the immense and intricate nature of ocean data as well as the necessity for higher granularity and richness in knowledge. To alleviate these issues, we introduce OceanGPT, the first-ever LLM in the ocean domain, which is expert in various ocean science tasks. We propose DoInstruct, a novel framework to automatically obtain a large volume of ocean domain instruction data, which generates instructions based on multi-agent collaboration. Additionally, we construct the first oceanography benchmark, OceanBench, to evaluate the capabilities of LLMs in the ocean domain. Though comprehensive experiments, OceanGPT not only shows a higher level of knowledge expertise for oceans science tasks but also gains preliminary embodied intelligence capabilities in ocean technology. Codes, data and checkpoints will soon be available at https://github.com/zjunlp/KnowLM.
翻译:海洋科学深入研究覆盖地球表面超过70%的海洋——这一生命与生物多样性的宝库,具有重大科学意义。近年来,大型语言模型(LLMs)的进步已彻底改变了科学研究的范式。尽管在其他领域取得了成功,但当前的大语言模型往往难以满足海洋学等领域专家的特定需求,LLMs在海洋科学中的潜力尚未得到充分探索。其根本原因可能在于海洋数据的庞大复杂特性,以及知识颗粒度和丰富性的更高要求。为解决这些问题,我们首次提出了海洋领域的大语言模型OceanGPT,该模型精通各类海洋科学任务。我们设计了创新框架DoInstruct,通过多智能体协作自动生成大量海洋领域指令数据。同时,我们构建了首个海洋学基准OceanBench,用于评估LLMs在海洋领域的能力。通过全面实验,OceanGPT不仅在海洋科学任务中展现出更高水平的知识专长,还在海洋技术领域初步获得了具身智能能力。相关代码、数据及模型参数即将在https://github.com/zjunlp/KnowLM 开放获取。