The automated program repair field has attracted substantial interest over the years, but despite significant research efforts, creating a system that works well for complex semantic bugs such as security vulnerabilities has proven difficult. A promising direction to solve this challenge is by leveraging large language models (LLMs), which are increasingly used to solve various programming tasks. In this paper, we investigate the effectiveness of LLMs for solving code-repair task. We show that the task is difficult as it requires the model to learn long-range code relationships, a task that inherently relies on extensive amounts of training data. At the same time, creating a large, clean dataset for complex program bugs and their corresponding fixes is non-trivial. We propose a technique to address these challenges with a new approach for querying and fine-tuning LLMs. The idea is to use program analysis to limit the LLM's attention mechanism on the portions of code needed to perform the fix, drastically reducing the amount of required training data. Concretely, for training and inference, rather than feeding the entire program to the LLM, we reduce its code to a much shorter snippet that contains the reported defect together with the necessary context - and use that instead. Our evaluation shows that this code reduction approach substantially improves available models such as GPT-4 using few-shot learning, as well as fine-tuning models. To train and evaluate our system, we created a comprehensive code fixing dataset by extensively labeling 156 bug patterns (including 40 security rules), requiring complex interprocedural dataflow to discover. Our best system with Mixtral-8x7B can remove more than 80% of the reported defects while exactly matching the human fix in between 10 and 50% of cases, outperforming baselines based on GPT-3.5 and GPT-4, or based on window-based models like TFix.
翻译:自动化程序修复领域多年来一直备受关注,但尽管投入了大量研究,构建一个能有效处理复杂语义错误(如安全漏洞)的系统仍具挑战性。解决这一难题的一个有前景方向是利用大语言模型(LLMs),该模型正被越来越多地用于解决各类编程任务。本文研究了LLMs在代码修复任务中的有效性。研究表明,该任务要求模型学习长距离代码依赖关系,本质上有赖于海量训练数据,因此难度较大。同时,为复杂程序错误及其对应修复构建大规模清晰数据集也非易事。我们提出了一种新技术,通过新型查询和微调LLMs的方法来应对这些挑战。其核心思路是利用程序分析技术,将LLM的注意力机制限制在修复所需的关键代码片段上,从而大幅减少所需训练数据量。具体而言,在训练和推理过程中,我们不再将完整程序输入LLM,而是将其代码缩减为包含缺陷报告及必要上下文的更短片段,并以此替代完整程序。实验评估表明,这种代码缩减方法能显著提升GPT-4等现有模型的少样本学习能力及微调效果。为训练和评估系统,我们构建了综合代码修复数据集,人工标注了156种缺陷模式(含40项安全规则),这些模式需要复杂的过程间数据流分析才能发现。采用Mixtral-8x7B的最佳系统可消除80%以上的报告缺陷,其中10%至50%的修复结果与人工修复完全一致,性能超越基于GPT-3.5、GPT-4的基线模型以及TFix等基于窗口的模型。