Building infrastructure-as-code (IaC) in cloud computing is a critical task, underpinning the reliability, scalability, and security of modern software systems. Despite the remarkable progress of large language models (LLMs) in software engineering -- demonstrated across many dedicated benchmarks -- their capabilities in developing IaC remain underexplored. Unlike existing IaC benchmarks that predominantly center on declarative paradigms such as Terraform and involve generating entire codebases from scratch, our benchmark reflects the incremental code edits common in enterprise development with imperative tools like the AWS CDK. We present SWE-InfraBench, a diverse evaluation dataset sourced from dozens of real-world IaC codebases that challenge LLMs to perform realistic code modifications in AWS CDK repositories. Each example requires models to implement changes to existing codebases based on natural language instructions, with success determined by passing provided test cases. These tasks demand sophisticated reasoning about cloud resource dependencies and implementation patterns beyond conventional code generation challenges. Our evaluation results reveal significant limitations in current LLMs showing that even state-of-the-art systems struggle with many tasks -- the best model, Sonnet 3.7, succeeds in only 34\% of cases, while specialized reasoning models like DeepSeek R1 achieve just 24% success. The SWE-InfraBench dataset is available at: https://www.kaggle.com/datasets/64e59070fd51c0278560b01eb5dc4f3c447d5268cdabe5a350d2969e4413fea5
翻译:构建云计算中的基础设施即代码(IaC)是一项关键任务,支撑着现代软件系统的可靠性、可扩展性和安全性。尽管大型语言模型(LLMs)在软件工程领域取得了显著进展(已在众多专门基准测试中得到证实),但它们在开发IaC方面的能力仍未得到充分探索。与现有主要关注声明式范式(如Terraform)且涉及从头生成完整代码库的IaC基准测试不同,我们的基准反映了企业开发中使用命令式工具(如AWS CDK)进行增量代码编辑的常见场景。我们提出了SWE-InfraBench,这是一个多样化的评估数据集,来源于数十个真实世界的IaC代码库,挑战LLMs在AWS CDK存储库中执行逼真的代码修改。每个示例要求模型根据自然语言指令对现有代码库实施变更,并通过提供的测试用例判定成功与否。这些任务要求对云资源依赖关系和实现模式进行复杂的推理,超越了传统的代码生成挑战。我们的评估结果揭示了当前LLMs的重大局限性,表明即使是最先进的系统在许多任务上也难以应对——最佳模型Sonnet 3.7仅成功处理了34%的案例,而专用推理模型如DeepSeek R1的成功率仅为24%。SWE-InfraBench数据集可在以下网址获取:https://www.kaggle.com/datasets/64e59070fd51c0278560b01eb5dc4f3c447d5268cdabe5a350d2969e4413fea5