Causal Inference offers a fundamental approach for advancing empirical software engineering (ESE) beyond traditional statistical association, enabling researchers to rigorously identify and quantify causal relationships in software experiments. This paper introduces CausalSE, a framework that operationalizes Judea Pearl's causal inference paradigm in ESE context. The paper focuses on Structural Causal Models (SCMs) to address the limitations of classical statistical methods in mitigating confounding bias. Through a case study using the Galeras dataset and propensity score matching, we demonstrate how CausalSE disentangles the effect of prompt engineering strategies on code generation outcomes in a popular LLM (i.e., GPT-3). The results reveal that while associational analyses can suggest improvements in certain interventions (e.g., more complex prompts), causal analysis often does not find a significant treatment effect, highlighting the risk of false positives when confounding is not addressed. By providing a tutorial-based methodology and a real-world case study, this work equips software researchers with practical tools to design, analyze, and interpret software experiments with methodological rigor, ultimately enabling more informed and actionable conclusions in both research and practice.
翻译:因果推断为超越传统统计关联推进经验软件工程(ESE)提供了一种基础方法,使研究人员能够严格识别和量化软件实验中的因果关系。本文介绍了CausalSE框架,该框架将Judea Pearl的因果推断范式应用于ESE场景。论文聚焦于结构因果模型(SCM),以解决经典统计方法在减轻混杂偏倚方面的局限性。通过使用Galeras数据集和倾向得分匹配的案例研究,我们展示了CausalSE如何分离提示工程策略对流行大语言模型(即GPT-3)代码生成结果的影响。结果表明,虽然关联分析可能提示某些干预措施(如更复杂的提示)能带来改进,但因果分析通常未发现显著的处理效应,这突显了在未解决混杂问题时出现假阳性结果的风险。通过提供基于教程的方法论和真实世界案例研究,本文为软件研究人员提供了实用的工具,使其能够以严谨的方法论设计、分析和解释软件实验结果,最终在研究与实践层面得出更明智且可操作的结论。