Bash scripts are critical for system administration, DevOps, and CI/CD, where code quality affects stability and security. However, LLM-generated scripts often lack reasoning and contain robustness flaws such as mishandled edge cases and unchecked failures. This is particularly critical in production environments where even minor errors can lead to service disruptions. We propose BashCoder-R1, a framework that jointly addresses both issues by treating explainability as a design goal. The pipeline has three stages. Continual Pre-training adapts to Bash syntax. Long Chain-of-Thought Supervised Fine-Tuning on expert-validated samples teaches risk-averse reasoning before code generation. Robustness-Aware Group Relative Policy Optimization optimizes a weighted reward for syntax correctness, robustness (verified by shellcheck), and format adherence. This staged design ensures that the model progressively acquires syntax knowledge, reasoning capability, and robust decision-making. On our BashBench benchmark (952 real-world tasks, 773 single-line and 179 multi-line), BashCoder-R1 achieves SyntaxPass of 100.00/94.97, RobustWarnRate of 4.01/16.47, RobustPass of 95.99/79.33, FuncRate of 93.01/93.85, and FullRate of 90.04/73.18 for single-line and multi-line tasks, respectively. These are relative FullRate improvements of 37.82 and 20.18 percent over the strongest baseline, DeepSeek-V3.2 (Reasoning). Human evaluation confirms its reasoning chains are highest in quality.


翻译:暂无翻译

0
下载
关闭预览

相关内容

Stabilizing Transformers for Reinforcement Learning
专知会员服务
61+阅读 · 2019年10月17日
【Github】GPT2-Chinese:中文的GPT2训练代码
AINLP
52+阅读 · 2019年8月23日
分布式并行架构Ray介绍
CreateAMind
10+阅读 · 2019年8月9日
近期语音类前沿论文
深度学习每日摘要
14+阅读 · 2019年3月17日
Focal Loss for Dense Object Detection
统计学习与视觉计算组
12+阅读 · 2018年3月15日
tensorflow LSTM + CTC实现端到端OCR
机器学习研究会
26+阅读 · 2017年11月16日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
分层反无人机系统发展新趋势
专知会员服务
7+阅读 · 9月3日
何为协作武器?
专知会员服务
10+阅读 · 9月1日
《理解认知战:超越信息》
专知会员服务
14+阅读 · 9月1日
美国战争部在GenAI.mil上推出OpenAI的ChatGPT Mil
专知会员服务
9+阅读 · 8月31日
人工智能赋能军事维护:重新定义国防战备
专知会员服务
5+阅读 · 8月31日
《美陆军野战手册(2026年):特种部队》
专知会员服务
9+阅读 · 8月31日
相关VIP内容
Stabilizing Transformers for Reinforcement Learning
专知会员服务
61+阅读 · 2019年10月17日
相关资讯
【Github】GPT2-Chinese:中文的GPT2训练代码
AINLP
52+阅读 · 2019年8月23日
分布式并行架构Ray介绍
CreateAMind
10+阅读 · 2019年8月9日
近期语音类前沿论文
深度学习每日摘要
14+阅读 · 2019年3月17日
Focal Loss for Dense Object Detection
统计学习与视觉计算组
12+阅读 · 2018年3月15日
tensorflow LSTM + CTC实现端到端OCR
机器学习研究会
26+阅读 · 2017年11月16日
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员