High-quality smart contract auditing datasets are crucial for evaluating security tools and advancing smart contract security research. Two major limitations of existing datasets are the manual-induced scalability bottleneck and the deficiency in data granularity and diversity. To address these limitations, we propose GiANT, an automated framework designed to curate smart contract auditing datasets by distilling vulnerability insights from real-world auditing reports. GiANT employs a divide-and-conquer strategy coupled with the Chain-of-Thought technique to extract structured vulnerability information from Code4rena reports, followed by an LLM-as-a-judge mechanism to perform rigorous quality assurance. To evaluate GiANT's effectiveness, we run it on 388 real-world audit reports and generate the GiAnt Corpus comprising 7,711 vulnerability findings across five severity levels. Manual assessment of the dataset demonstrates exceptional reliability in information extraction, achieving a mean quality score of $4.76\pm0.37$ (out of 5) with inter-rater agreement $κ$ of 0.88. We further validate the practicality of our dataset by benchmarking 4 state-of-the-art LLMs on vulnerability detection, code summarization, mitigation recommendation, and automated gas optimization tasks, to establish performance baselines, thereby providing a valuable data foundation for future research in automated smart contract auditing.
翻译:高质量的智能合约审计数据集对于评估安全工具和推动智能合约安全研究至关重要。现有数据集的两大主要缺陷是:人工驱动的可扩展性瓶颈,以及数据粒度与多样性的不足。为克服这些局限,我们提出了GiANT——一个通过从真实世界审计报告中提炼漏洞洞见来整理智能合约审计数据集的自动化框架。GiANT采用分治策略,结合思维链技术从Code4rena报告中提取结构化漏洞信息,并引入“大语言模型作为法官”机制进行严格的质量保证。为评估GiANT的有效性,我们针对388份真实审计报告运行该框架,生成了包含7,711条漏洞发现(涵盖五个严重级别)的GiAnt语料库。对数据集的人工评估表明,信息提取具有卓越的可靠性,平均质量得分达到4.76±0.37(满分5分),评分者间一致性系数κ为0.88。我们进一步通过基准测试4个最先进的大语言模型,在漏洞检测、代码摘要、缓解建议及自动Gas优化任务上验证了该数据集的实用性,建立了性能基线,从而为未来自动化智能合约审计研究提供了宝贵的数据基础。