Ensuring correctness is a pivotal aspect of software engineering. Among the various strategies available, software verification offers a definitive assurance of correctness. Nevertheless, writing verification proofs is resource-intensive and manpower-consuming, and there is a great need to automate this process. We introduce Selene in this paper, which is the first project-level automated proof benchmark constructed based on the real-world industrial-level project of the seL4 operating system microkernel. Selene provides a comprehensive framework for end-to-end evaluation and a lightweight verification environment. Our experimental results with advanced LLMs, such as GPT-3.5-turbo and GPT-4, highlight the capabilities of large language models (LLMs) in the domain of automated proof generation. Additionally, our further proposed augmentations indicate that the challenges presented by Selene can be mitigated in future research endeavors.
翻译:确保证正确性是软件工程中的关键环节。在现有多种策略中,软件验证能够提供确定性的正确性保障。然而,编写验证证明既消耗大量计算资源又需要人力投入,因此自动化这一过程的需求极为迫切。本文提出Selene——首个基于真实工业级项目seL4操作系统微内核构建的项目级自动化证明基准测试框架。Selene提供全面的端到端评估框架和轻量级验证环境。我们基于GPT-3.5-turbo和GPT-4等先进大语言模型的实验结果,揭示了大型语言模型在自动化证明生成领域的潜力。此外,我们进一步提出的增强方案表明,Selene所呈现的挑战可在未来研究工作中得到有效缓解。