We introduce TeleQnA, the first benchmark dataset designed to evaluate the knowledge of Large Language Models (LLMs) in telecommunications. Comprising 10,000 questions and answers, this dataset draws from diverse sources, including standards and research articles. This paper outlines the automated question generation framework responsible for creating this dataset, along with how human input was integrated at various stages to ensure the quality of the questions. Afterwards, using the provided dataset, an evaluation is conducted to assess the capabilities of LLMs, including GPT-3.5 and GPT-4. The results highlight that these models struggle with complex standards related questions but exhibit proficiency in addressing general telecom-related inquiries. Additionally, our results showcase how incorporating telecom knowledge context significantly enhances their performance, thus shedding light on the need for a specialized telecom foundation model. Finally, the dataset is shared with active telecom professionals, whose performance is subsequently benchmarked against that of the LLMs. The findings illustrate that LLMs can rival the performance of active professionals in telecom knowledge, thanks to their capacity to process vast amounts of information, underscoring the potential of LLMs within this domain. The dataset has been made publicly accessible on GitHub.
翻译:我们提出了TeleQnA,这是首个专为评估大语言模型(LLMs)在电信领域知识能力而设计的基准数据集。该数据集包含10,000个问答对,数据来源涵盖标准文件和研究论文等多种资源。本文阐述了用于生成该数据集的自动化问答生成框架,以及如何在不同阶段融入人工输入以确保问题质量。随后,利用该数据集对GPT-3.5和GPT-4等LLMs的能力进行了评估。结果表明,这些模型在处理复杂的标准相关问题时表现不佳,但在回答一般性电信领域问题时展现出较强能力。此外,我们的结果揭示了加入电信知识上下文能显著提升模型性能,从而突显了开发专用电信基础模型的必要性。最后,我们将数据集分享给在职电信专业人员,并将其答题表现与LLMs进行对比基准测试。研究发现,凭借处理海量信息的能力,LLMs在电信知识方面可与在职专业人员相媲美,这充分展现了LLMs在该领域的应用潜力。该数据集已在GitHub上公开发布。