Large Language Models (LLMs) are widely used in software engineering (SE) research and practice, yet their non-determinism, opaque training data, and rapidly evolving models threaten the reproducibility and replicability of empirical studies. We address this challenge through a collaborative effort of 22 researchers, presenting a taxonomy of seven study types that organizes how LLMs are used in SE research, together with eight guidelines for designing and reporting such studies. Each guideline distinguishes requirements (must) from recommendations (should) and is contextualized by the study types it applies to. Our guidelines recommend that researchers: (1) declare LLM usage and role; (2) report model versions, configurations, and customizations; (3) document the system and prompt design beyond the model; (4) report session traces, i.e., interaction logs and runtime traces; (5) use suitable baselines, benchmarks, and metrics; (6) include an open LLM as a baseline; (7) validate LLM outputs against human judgment; and (8) articulate limitations and mitigations. We complement the guidelines with an applicability matrix mapping guidelines to study types and a reporting checklist for authors and reviewers. We maintain the study types and guidelines online as a living resource for the community to use and shape (llm-guidelines$.$org).
翻译:大语言模型(LLMs)在软件工程(SE)研究和实践中被广泛应用,但其非确定性、不透明的训练数据以及快速演化的模型威胁着实证研究的可重现性和可重复性。我们通过22名研究人员的协作努力应对这一挑战,提出了一种包含七种研究类型的分类体系,以组织LLMs在软件工程研究中的使用方式,并配套制定了八条关于此类研究设计与报告的建议。每条建议区分了强制性要求(必须)与推荐性实践(建议),并根据其适用的研究类型进行了情境化说明。我们的建议要求研究人员:(1)声明LLM的使用方式及角色;(2)报告模型版本、配置与定制化信息;(3)在模型之外记录系统与提示设计;(4)报告会话轨迹(即交互日志与运行时轨迹);(5)使用合适的基线、基准与评估指标;(6)将开源LLM作为基线;(7)通过人工判断验证LLM输出结果;(8)阐明局限性及缓解措施。我们以适用性矩阵(将建议映射至研究类型)及供作者与审稿人使用的报告检查清单作为补充。我们在llm-guidelines$.$org上持续维护这些研究类型与建议,作为供社区使用与完善的动态资源。