Prompting has become one of the main approaches to leverage emergent capabilities of Large Language Models [Brown et al. NeurIPS 2020, Wei et al. TMLR 2022, Wei et al. NeurIPS 2022]. During the last year, researchers and practitioners have been playing with prompts to see how to make the most of LLMs. By homogeneously dissecting 80 papers, we investigate in deep how software testing and verification research communities have been abstractly architecting their LLM-enabled solutions. More precisely, first, we want to validate whether downstream tasks are an adequate concept to convey the blueprint of prompt-based solutions. We also aim at identifying number and nature of such tasks in solutions. For such goal, we develop a novel downstream task taxonomy that enables pinpointing some engineering patterns in a rather varied spectrum of Software Engineering problems that encompasses testing, fuzzing, debugging, vulnerability detection, static analysis and program verification approaches.
翻译:提示已成为利用大语言模型涌现能力的主要方法之一 [Brown 等, NeurIPS 2020; Wei 等, TMLR 2022; Wei 等, NeurIPS 2022]。过去一年间,研究人员和从业者不断探索提示技术,以期最大化发挥大语言模型的作用。通过同质化剖析80篇论文,我们深入探究了软件测试与验证研究社区如何抽象架构其基于大语言模型的解决方案。具体而言,我们首先验证下游任务是否足以传达基于提示的解决方案蓝图,其次旨在识别此类解决方案中任务的数量及其本质。为此,我们提出一种新颖的下游任务分类方法,可精确定位涵盖测试、模糊测试、调试、漏洞检测、静态分析与程序验证方法等广泛软件工程问题中的若干工程模式。