Impact of Code Language Models on Automated Program Repair - 专知论文

会员服务 ·

0

基准测试 · 代码 · 基准 · 软件可靠性 · 语言模型 ·

2023 年 4 月 16 日

Impact of Code Language Models on Automated Program Repair

翻译：代码语言模型对自动程序修复的影响

Nan Jiang,Kevin Liu,Thibaud Lutellier,Lin Tan

from arxiv, This paper is accepted by 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE)

Automated program repair (APR) aims to help developers improve software reliability by generating patches for buggy programs. Although many code language models (CLM) are developed and effective in many software tasks such as code completion, there has been little comprehensive, in-depth work to evaluate CLMs' fixing capabilities and to fine-tune CLMs for the APR task. Firstly, this work is the first to evaluate ten CLMs on four APR benchmarks, which shows that surprisingly, the best CLM, as is, fixes 72% more bugs than the state-of-the-art deep-learning (DL)-based APR techniques. Secondly, one of the four APR benchmarks was created by us in this paper to avoid data leaking for a fair evaluation. Thirdly, it is the first work to fine-tune CLMs with APR training data, which shows that fine-tuning brings 31%-1,267% improvement to CLMs and enables them to fix 46%-164% more bugs than existing DL-based APR techniques. Fourthly, this work studies the impact of buggy lines, showing that CLMs, as is, cannot make good use of the buggy lines to fix bugs, yet fine-tuned CLMs could potentially over-rely on buggy lines. Lastly, this work analyzes the size, time, and memory efficiency of different CLMs. This work shows promising directions for the APR domain, such as fine-tuning CLMs with APR-specific designs, and also raises awareness of fair and comprehensive evaluations of CLMs and calls for more transparent reporting of open-source repositories used in the pre-training data to address the data leaking problem.

翻译：自动程序修复（APR）旨在通过为缺陷程序生成补丁来帮助开发者提升软件可靠性。尽管许多代码语言模型（CLM）在代码补全等软件任务中表现出色，但缺乏全面深入的评估工作来检验其修复能力，也鲜有研究针对APR任务对CLM进行微调。首先，本研究首次在四个APR基准上评估了十种CLM，结果出人意料地表明：未经调整的最佳CLM修复的缺陷数量比最先进的基于深度学习的APR技术多72%。其次，四个APR基准中有一个由本文创建，旨在避免数据泄露以实现公平评估。第三，本研究首次使用APR训练数据对CLM进行微调，结果发现微调可为CLM带来31%至1,267%的性能提升，使其修复的缺陷比现有基于深度学习的APR技术多46%至164%。第四，本研究分析了缺陷行的作用，发现未经微调的CLM难以有效利用缺陷行修复缺陷，而微调后的CLM可能过度依赖缺陷行。最后，本研究分析了不同CLM的规模、时间效率和内存效率。这项工作为APR领域指明了有前景的方向（例如针对APR特定设计微调CLM），同时呼吁对CLM进行公平全面的评估，并要求更透明地报告预训练数据中使用的开源仓库，以解决数据泄露问题。

0

相关内容

基准测试

基准测试是指通过设计科学的测试方法、测试工具和测试系统，实现对一类测试对象的某项性能指标进行定量的和可对比的测试。

【Manning新书】自动机器学习实战，Automated Machine Learning in Action

【Manning新书】自动机器学习实战，Automated Machine Learning in Action

专知会员服务

95+阅读 · 2022年4月8日

【干货书】深度学习合成数据，354页pdf，Synthetic Data for Deep Learning

【干货书】深度学习合成数据，354页pdf，Synthetic Data for Deep Learning

专知会员服务

105+阅读 · 2022年2月10日

【KDD2020-Tutorial】自动推荐系统，Automated Recommendation System

【KDD2020-Tutorial】自动推荐系统，Automated Recommendation System

专知会员服务

53+阅读 · 2020年8月25日

【微软】利用知识图谱提高抽象摘要的事实正确性，Boosting Factual Correctness

专知会员服务

18+阅读 · 2020年3月23日

50+篇《神经架构搜索NAS》2020论文合集

专知会员服务

61+阅读 · 2020年3月19日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

167+阅读 · 2020年3月18日

八篇 ICCV 2019 【图神经网络（GNN）+CV】相关论文

八篇 ICCV 2019 【图神经网络（GNN）+CV】相关论文

专知会员服务

30+阅读 · 2020年1月10日

【论文】把人类从学习应用中带出来：自动机器学习综述（Taking the Human out of Learning Applications: A Survey on Automated Machine Learning）

【论文】把人类从学习应用中带出来：自动机器学习综述（Taking the Human out of Learning Applications: A Survey on Automated Machine Learning）

专知会员服务

12+阅读 · 2019年12月20日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

论文浅尝 | Language Models (Mostly) Know What They Know

论文浅尝 | Language Models (Mostly) Know What They Know

开放知识图谱

2+阅读 · 2022年11月18日

“全职做开源 6 个月，我真的不后悔”

“全职做开源 6 个月，我真的不后悔”

CSDN

0+阅读 · 2022年9月21日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

META微软等最新ACL2022教程《非自回归序列生成》，168页ppt

META微软等最新ACL2022教程《非自回归序列生成》，168页ppt

专知

2+阅读 · 2022年6月3日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知

133+阅读 · 2020年3月18日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

44+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

Rho/ROCK信号通路介导的侵入性死亡（Entosis）在去势抵抗性前列腺癌中的作用及其机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

MDSCs调控piRNA介导DNA甲基化参与骨髓瘤干细胞形成及耐药的分子机制

国家自然科学基金

0+阅读 · 2015年12月31日

长链非编码RNA CAR intergenic 10在细胞衰老中的作用和机制

国家自然科学基金

1+阅读 · 2013年12月31日

SAH在ApoE-/-小鼠动脉粥样硬化形成中的作用机制及甜菜碱干预研究

国家自然科学基金

0+阅读 · 2013年12月31日

多孔POSS/PDMS分子内杂化膜的制备及其渗透汽化优先透醇性能研究

国家自然科学基金

0+阅读 · 2012年12月31日

miR34c重启衰老清除急性髓系白血病干细胞与机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

SCN5A突变(D772N和A1656V)致重叠型室性心律失常机制的研究

国家自然科学基金

0+阅读 · 2011年12月31日

含缺陷桩的灌注桩基础竖向承载性状研究

国家自然科学基金

0+阅读 · 2009年12月31日

柔性铜铟镓硒太阳电池异质结的调控及其对光伏性能的影响

国家自然科学基金

0+阅读 · 2009年12月31日

深基坑卸载后的坑底地基与既有工程桩的受力变形性状研究

国家自然科学基金

0+阅读 · 2009年12月31日

Multilingual Conceptual Coverage in Text-to-Image Models

Arxiv

0+阅读 · 2023年6月2日

Generation of Probabilistic Synthetic Data for Serious Games: A Case Study on Cyberbullying

Arxiv

0+阅读 · 2023年6月2日

The Hidden Language of Diffusion Models

Arxiv

0+阅读 · 2023年6月1日

ReFACT: Updating Text-to-Image Models by Editing the Text Encoder

Arxiv

0+阅读 · 2023年6月1日

Better Context Makes Better Code Language Models: A Case Study on Function Call Argument Completion

Arxiv

0+阅读 · 2023年6月1日

Red Teaming Language Model Detectors with Language Models

Arxiv

0+阅读 · 2023年5月31日

A Survey of Knowledge-Enhanced Pre-trained Language Models

Arxiv

18+阅读 · 2022年11月17日

QA-GNN: Reasoning with Language Models and Knowledge Graphs for Question Answering

Arxiv

20+阅读 · 2021年5月27日

Machine Reading Comprehension: The Role of Contextualized Language Models and Beyond

Arxiv

15+阅读 · 2020年5月13日

Taking Human out of Learning Applications: A Survey on Automated Machine Learning

Taking Human out of Learning Applications: A Survey on Automated Machine Learning

Arxiv

14+阅读 · 2019年1月17日

VIP会员

文章信息

相关主题

软件可靠性

最新内容

ICML 2026 | FR3D：解耦自车运动的未来动态三维重建世界模型

ICML 2026 | FR3D：解耦自车运动的未来动态三维重建世界模型

专知会员服务

1+阅读 · 今天14:49

【伯克利博士论文】迈向可扩展与自我演进的大语言模型智能体

【伯克利博士论文】迈向可扩展与自我演进的大语言模型智能体

专知会员服务

1+阅读 · 今天14:47

学习数据的几何：形状空间分析数学综述

学习数据的几何：形状空间分析数学综述

专知会员服务

1+阅读 · 今天14:45

《现代防空系统综述：架构、传感器、拦截器及新兴威胁环境对基础设施受限防御环境的影响》2026最新长综述

《现代防空系统综述：架构、传感器、拦截器及新兴威胁环境对基础设施受限防御环境的影响》2026最新长综述

专知会员服务

2+阅读 · 今天14:22

定向能反无人机系统最新发展动态

定向能反无人机系统最新发展动态

专知会员服务

3+阅读 · 今天13:50

从燃煤战舰到算法战争：水面指挥的永恒要求

从燃煤战舰到算法战争：水面指挥的永恒要求

专知会员服务

2+阅读 · 今天13:33

《短程弹道再入飞行器拦截时间中的一项异常现象》

《短程弹道再入飞行器拦截时间中的一项异常现象》

专知会员服务

2+阅读 · 今天13:30

《基于回归方法与任务上下文的对抗环境动态战术网络报文优先级排序》

《基于回归方法与任务上下文的对抗环境动态战术网络报文优先级排序》

专知会员服务

2+阅读 · 今天13:28

美智库《战术级指挥控制的迫切要求：构建弹性机动式指挥控制网络》报告

美智库《战术级指挥控制的迫切要求：构建弹性机动式指挥控制网络》报告

专知会员服务

2+阅读 · 今天13:13

《韩国国防政策与军备出口：韩国安全与国防政策如何塑造其国防工业与军备出口格局》最新100页报告

《韩国国防政策与军备出口：韩国安全与国防政策如何塑造其国防工业与军备出口格局》最新100页报告

专知会员服务

1+阅读 · 今天13:10

ICML 2026 | VOTP：用视频基础模型与最优传输，让离线偏好强化学习只需少量反馈

ICML 2026 | VOTP：用视频基础模型与最优传输，让离线偏好强化学习只需少量反馈

专知会员服务

5+阅读 · 6月16日

多模态代码智能综述：从视觉输入到可执行代码系统

多模态代码智能综述：从视觉输入到可执行代码系统

专知会员服务

7+阅读 · 6月16日

美国马六甲“三重网”概念：安全网、威慑网与杀伤网

美国马六甲“三重网”概念：安全网、威慑网与杀伤网

专知会员服务

5+阅读 · 6月16日

《面向导弹有效发射时机的监督机器学习方法：基于超视距空战仿真》

《面向导弹有效发射时机的监督机器学习方法：基于超视距空战仿真》

专知会员服务

5+阅读 · 6月16日

《通用大语言模型：无人机指挥与控制接口》最新40页

《通用大语言模型：无人机指挥与控制接口》最新40页

专知会员服务

15+阅读 · 6月16日

相关VIP内容

【Manning新书】自动机器学习实战，Automated Machine Learning in Action

【Manning新书】自动机器学习实战，Automated Machine Learning in Action

专知会员服务

95+阅读 · 2022年4月8日

【干货书】深度学习合成数据，354页pdf，Synthetic Data for Deep Learning

【干货书】深度学习合成数据，354页pdf，Synthetic Data for Deep Learning

专知会员服务

105+阅读 · 2022年2月10日

【KDD2020-Tutorial】自动推荐系统，Automated Recommendation System

【KDD2020-Tutorial】自动推荐系统，Automated Recommendation System

专知会员服务

53+阅读 · 2020年8月25日

【微软】利用知识图谱提高抽象摘要的事实正确性，Boosting Factual Correctness

专知会员服务

18+阅读 · 2020年3月23日

50+篇《神经架构搜索NAS》2020论文合集

专知会员服务

61+阅读 · 2020年3月19日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

167+阅读 · 2020年3月18日

八篇 ICCV 2019 【图神经网络（GNN）+CV】相关论文

八篇 ICCV 2019 【图神经网络（GNN）+CV】相关论文

专知会员服务

30+阅读 · 2020年1月10日

【论文】把人类从学习应用中带出来：自动机器学习综述（Taking the Human out of Learning Applications: A Survey on Automated Machine Learning）

【论文】把人类从学习应用中带出来：自动机器学习综述（Taking the Human out of Learning Applications: A Survey on Automated Machine Learning）

专知会员服务

12+阅读 · 2019年12月20日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

热门VIP内容

开通专知VIP会员享更多权益服务

【伯克利博士论文】迈向可扩展与自我演进的大语言模型智能体

《现代防空系统综述：架构、传感器、拦截器及新兴威胁环境对基础设施受限防御环境的影响》2026最新长综述

ICML 2026 | FR3D：解耦自车运动的未来动态三维重建世界模型

学习数据的几何：形状空间分析数学综述

相关资讯

论文浅尝 | Language Models (Mostly) Know What They Know

论文浅尝 | Language Models (Mostly) Know What They Know

开放知识图谱

2+阅读 · 2022年11月18日

“全职做开源 6 个月，我真的不后悔”

“全职做开源 6 个月，我真的不后悔”

CSDN

0+阅读 · 2022年9月21日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

META微软等最新ACL2022教程《非自回归序列生成》，168页ppt

META微软等最新ACL2022教程《非自回归序列生成》，168页ppt

专知

2+阅读 · 2022年6月3日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知

133+阅读 · 2020年3月18日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

44+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

相关论文

Multilingual Conceptual Coverage in Text-to-Image Models

Arxiv

0+阅读 · 2023年6月2日

Generation of Probabilistic Synthetic Data for Serious Games: A Case Study on Cyberbullying

Arxiv

0+阅读 · 2023年6月2日

The Hidden Language of Diffusion Models

Arxiv

0+阅读 · 2023年6月1日

ReFACT: Updating Text-to-Image Models by Editing the Text Encoder

Arxiv

0+阅读 · 2023年6月1日

Better Context Makes Better Code Language Models: A Case Study on Function Call Argument Completion

Arxiv

0+阅读 · 2023年6月1日

Red Teaming Language Model Detectors with Language Models

Arxiv

0+阅读 · 2023年5月31日

A Survey of Knowledge-Enhanced Pre-trained Language Models

Arxiv

18+阅读 · 2022年11月17日

QA-GNN: Reasoning with Language Models and Knowledge Graphs for Question Answering

Arxiv

20+阅读 · 2021年5月27日

Machine Reading Comprehension: The Role of Contextualized Language Models and Beyond

Arxiv

15+阅读 · 2020年5月13日

Taking Human out of Learning Applications: A Survey on Automated Machine Learning

Taking Human out of Learning Applications: A Survey on Automated Machine Learning

Arxiv

14+阅读 · 2019年1月17日

相关基金

Rho/ROCK信号通路介导的侵入性死亡（Entosis）在去势抵抗性前列腺癌中的作用及其机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

MDSCs调控piRNA介导DNA甲基化参与骨髓瘤干细胞形成及耐药的分子机制

国家自然科学基金

0+阅读 · 2015年12月31日

长链非编码RNA CAR intergenic 10在细胞衰老中的作用和机制

国家自然科学基金

1+阅读 · 2013年12月31日

SAH在ApoE-/-小鼠动脉粥样硬化形成中的作用机制及甜菜碱干预研究

国家自然科学基金

0+阅读 · 2013年12月31日

多孔POSS/PDMS分子内杂化膜的制备及其渗透汽化优先透醇性能研究

国家自然科学基金

0+阅读 · 2012年12月31日

miR34c重启衰老清除急性髓系白血病干细胞与机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

SCN5A突变(D772N和A1656V)致重叠型室性心律失常机制的研究

国家自然科学基金

0+阅读 · 2011年12月31日

含缺陷桩的灌注桩基础竖向承载性状研究

国家自然科学基金

0+阅读 · 2009年12月31日

柔性铜铟镓硒太阳电池异质结的调控及其对光伏性能的影响

国家自然科学基金

0+阅读 · 2009年12月31日

深基坑卸载后的坑底地基与既有工程桩的受力变形性状研究

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员