Recently, various illustrative examples have shown the impressive ability of generative large language models (LLMs) to perform NLP related tasks. ChatGPT undoubtedly is the most representative model. We empirically evaluate ChatGPT's performance on requirements information retrieval (IR) tasks to derive insights into designing or developing more effective requirements retrieval methods or tools based on generative LLMs. We design an evaluation framework considering four different combinations of two popular IR tasks and two common artifact types. Under zero-shot setting, evaluation results reveal ChatGPT's promising ability to retrieve requirements relevant information (high recall) and limited ability to retrieve more specific requirements information (low precision). Our evaluation of ChatGPT on requirements IR under zero-shot setting provides preliminary evidence for designing or developing more effective requirements IR methods or tools based on LLMs.
翻译:近期,各类示例展示了生成式大语言模型(LLMs)在自然语言处理任务上的卓越能力,ChatGPT无疑是最具代表性的模型。我们通过实证评估ChatGPT在需求信息检索(IR)任务上的表现,为基于生成式LLMs设计或开发更高效的需求检索方法或工具提供洞见。我们设计了一个评估框架,涵盖两种常见IR任务与两种常见工件类型的四种不同组合。在零样本设置下,评估结果显示ChatGPT在检索需求相关信息方面展现出潜力(高召回率),但在获取更具体需求信息方面能力有限(低精确率)。我们对ChatGPT在零样本设置下进行需求IR的评估,为基于LLMs设计或开发更高效的需求IR方法或工具提供了初步依据。