High-quality smart contract auditing datasets are crucial for evaluating security tools and advancing smart contract security research. Two major limitations of existing datasets are the manual-induced scalability bottleneck and the deficiency in data granularity and diversity. To address these limitations, we propose GiANT, an automated framework designed to curate smart contract auditing datasets by distilling vulnerability insights from real-world auditing reports. GiANT employs a divide-and-conquer strategy coupled with the Chain-of-Thought technique to extract structured vulnerability information from Code4rena reports, followed by an LLM-as-a-judge mechanism to perform rigorous quality assurance. To evaluate GiANT's effectiveness, we run it on 388 real-world audit reports and generate the GiAnt Corpus comprising 7,711 vulnerability findings across five severity levels. Manual assessment of the dataset demonstrates exceptional reliability in information extraction, achieving a mean quality score of $4.76\pm0.37$ (out of 5) with inter-rater agreement $κ$ of 0.88. We further validate the practicality of our dataset by benchmarking 4 state-of-the-art LLMs on vulnerability detection, code summarization, mitigation recommendation, and automated gas optimization tasks, to establish performance baselines, thereby providing a valuable data foundation for future research in automated smart contract auditing.


翻译:高质量的智能合约审计数据集对于评估安全工具和推动智能合约安全研究至关重要。现有数据集的两大主要缺陷是:人工驱动的可扩展性瓶颈,以及数据粒度与多样性的不足。为克服这些局限,我们提出了GiANT——一个通过从真实世界审计报告中提炼漏洞洞见来整理智能合约审计数据集的自动化框架。GiANT采用分治策略,结合思维链技术从Code4rena报告中提取结构化漏洞信息,并引入“大语言模型作为法官”机制进行严格的质量保证。为评估GiANT的有效性,我们针对388份真实审计报告运行该框架,生成了包含7,711条漏洞发现(涵盖五个严重级别)的GiAnt语料库。对数据集的人工评估表明,信息提取具有卓越的可靠性,平均质量得分达到4.76±0.37(满分5分),评分者间一致性系数κ为0.88。我们进一步通过基准测试4个最先进的大语言模型,在漏洞检测、代码摘要、缓解建议及自动Gas优化任务上验证了该数据集的实用性,建立了性能基线,从而为未来自动化智能合约审计研究提供了宝贵的数据基础。

0
下载
关闭预览

相关内容

Google《AI智能体企业应用手册报告》,46页pdf
专知会员服务
50+阅读 · 2025年12月29日
美智库最新报告:小数据人工智能潜力不可估量,39页pdf
专知会员服务
77+阅读 · 2021年11月18日
《人工智能安全框架(2020年)》白皮书,68页pdf
专知会员服务
167+阅读 · 2021年1月9日
《人工智能安全测评白皮书》,99页pdf
专知
36+阅读 · 2022年2月26日
智能合约的形式化验证方法研究综述
专知
16+阅读 · 2021年5月8日
国家自然科学基金
1+阅读 · 2017年12月31日
国家自然科学基金
18+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
8+阅读 · 2015年12月31日
国家自然科学基金
8+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
Arxiv
0+阅读 · 5月27日
Arxiv
0+阅读 · 5月16日
VIP会员
最新内容
印度精确打击与指挥架构的断层
专知会员服务
4+阅读 · 7月20日
美空军AI完成F-16战斗机自主空战历史性试飞
专知会员服务
6+阅读 · 7月20日
深入Project Maven:为何人工智能在战场上依然失灵
锻造未来士兵:外骨骼、基因工程与赛博格
专知会员服务
7+阅读 · 7月19日
《无人机蜂群通信技术研究》50页
专知会员服务
10+阅读 · 7月19日
相关VIP内容
相关基金
国家自然科学基金
1+阅读 · 2017年12月31日
国家自然科学基金
18+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
8+阅读 · 2015年12月31日
国家自然科学基金
8+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员