Large Language Models (LLMs) have the potential to revolutionize the Sixth Generation (6G) communication networks. However, current mainstream LLMs generally lack the specialized knowledge in telecom domain. In this paper, for the first time, we propose a pipeline to adapt any general purpose LLMs to a telecom-specific LLMs. We collect and build telecom-specific pre-train dataset, instruction dataset, preference dataset to perform continual pre-training, instruct tuning and alignment tuning respectively. Besides, due to the lack of widely accepted evaluation benchmarks in telecom domain, we extend existing evaluation benchmarks and proposed three new benchmarks, namely, Telecom Math Modeling, Telecom Open QnA and Telecom Code Tasks. These new benchmarks provide a holistic evaluation of the capabilities of LLMs including math modeling, Open-Ended question answering, code generation, infilling, summarization and analysis in telecom domain. Our fine-tuned LLM TelecomGPT outperforms state of the art (SOTA) LLMs including GPT-4, Llama-3 and Mistral in Telecom Math Modeling benchmark significantly and achieve comparable performance in various evaluation benchmarks such as TeleQnA, 3GPP technical documents classification, telecom code summary and generation and infilling.
翻译:大语言模型(LLMs)具有革新第六代(6G)通信网络的潜力。然而,当前主流大语言模型普遍缺乏电信领域的专业知识。本文首次提出一种流程,可将任何通用大语言模型适配为电信领域专用大语言模型。我们收集并构建了电信领域专用的预训练数据集、指令数据集和偏好数据集,分别用于执行持续预训练、指令微调和对齐微调。此外,由于电信领域缺乏广泛接受的评估基准,我们扩展了现有评估基准,并提出了三个新基准,即电信数学建模、电信开放问答和电信代码任务。这些新基准为大语言模型在电信领域的数学建模、开放式问答、代码生成、填充、摘要和分析等能力提供了全面评估。我们微调后的大语言模型TelecomGPT在电信数学建模基准上显著优于包括GPT-4、Llama-3和Mistral在内的最先进大语言模型,并在TeleQnA、3GPP技术文档分类、电信代码摘要与生成及填充等多种评估基准中取得了可比的性能。