We present RM-RF, a lightweight reward model for run-free evaluation of automatically generated unit tests. Instead of repeatedly compiling and executing candidate tests, RM-RF predicts - from source and test code alone - three execution-derived signals: (1) whether the augmented test suite compiles and runs successfully, (2) whether the generated test cases increase code coverage, and (3) whether the generated test cases improve the mutation kill rate. To train and evaluate RM-RF we assemble a multilingual dataset (Java, Python, Go) of focal files, test files, and candidate test additions labeled by an execution-based pipeline, and we release an associated dataset and methodology for comparative evaluation. We tested multiple model families and tuning regimes (zero-shot, full fine-tuning, and PEFT via LoRA), achieving an average F1 of 0.69 across the three targets. Compared to conventional compile-and-run instruments, RM-RF provides substantially lower latency and infrastructure cost while delivering competitive predictive fidelity, enabling fast, scalable feedback for large-scale test generation and RL-based code optimization.


翻译:本文提出RM-RF,一种用于自动生成单元测试的免运行评估的轻量级奖励模型。该方法无需重复编译和执行候选测试,仅通过源代码和测试代码即可预测三个执行衍生信号:(1)增强后的测试套件能否成功编译运行,(2)生成的测试用例是否提高了代码覆盖率,(3)生成的测试用例是否改善了变异杀死率。为训练和评估RM-RF,我们构建了一个多语言(Java、Python、Go)数据集,包含核心文件、测试文件以及通过基于执行的流水线标注的候选测试增补,并发布了用于对比评估的配套数据集与方法论。我们测试了多种模型族与调优机制(零样本、全微调及基于LoRA的参数高效微调),在三个预测目标上平均F1值达到0.69。与传统编译运行工具相比,RM-RF在保持竞争力的预测保真度的同时,显著降低了延迟与基础设施成本,从而为大规模测试生成和基于强化学习的代码优化提供快速、可扩展的反馈机制。

0
下载
关闭预览

相关内容

代码(Code)是专知网的一个重要知识资料文档板块,旨在整理收录论文源代码、复现代码,经典工程代码等,便于用户查阅下载使用。
预训练语言模型fine-tuning近期进展概述
专知会员服务
40+阅读 · 2021年4月9日
深度学习在CTR预估中的应用 | CTR深度模型大盘点
PaperWeekly
15+阅读 · 2018年4月11日
深度学习目标检测模型全面综述:Faster R-CNN、R-FCN和SSD
深度学习世界
10+阅读 · 2017年9月18日
基于机器学习的KPI自动化异常检测系统
运维帮
13+阅读 · 2017年8月16日
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Arxiv
0+阅读 · 2月14日
VIP会员
最新内容
美空军如何将人工智能从战场部署至后方机关
专知会员服务
1+阅读 · 14分钟前
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
2+阅读 · 今天2:57
《英国防部:未来空战系统数字化战略》33页
专知会员服务
2+阅读 · 今天2:35
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
3+阅读 · 今天2:25
“史诗怒火”行动:现代多域作战的重要节点
专知会员服务
8+阅读 · 7月30日
《下一代无线网络中的多无人机通信资源管理》
相关VIP内容
预训练语言模型fine-tuning近期进展概述
专知会员服务
40+阅读 · 2021年4月9日
相关基金
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员