AssertFlip：通过反转LLM生成的通过测试来复现缺陷 (AssertFlip: Reproducing Bugs via Inversion of LLM-Generated Passing Tests)

Bug reproduction is critical in the software debugging and repair process, yet the majority of bugs in open-source and industrial settings lack executable tests to reproduce them at the time they are reported, making diagnosis and resolution more difficult and time-consuming. To address this challenge, we introduce AssertFlip, a novel technique for automatically generating Bug Reproducible Tests (BRTs) using large language models (LLMs). Unlike existing methods that attempt direct generation of failing tests, AssertFlip first generates passing tests on the buggy behaviour and then inverts these tests to fail when the bug is present. We hypothesize that LLMs are better at writing passing tests than ones that crash or fail on purpose. Our results show that AssertFlip outperforms all known techniques in the leaderboard of SWT-Bench, a benchmark curated for BRTs. Specifically, AssertFlip achieves a fail-to-pass success rate of 43.6% on the SWT-Bench-Verified subset.

翻译：缺陷复现在软件调试与修复过程中至关重要，然而在开源和工业环境中，大多数缺陷在被报告时缺乏可执行的测试用例来复现，这使得诊断和解决过程更加困难且耗时。为应对这一挑战，我们提出了AssertFlip，一种利用大语言模型自动生成缺陷可复现测试的新技术。与现有尝试直接生成失败测试的方法不同，AssertFlip首先生成针对缺陷行为的通过测试，随后将这些测试反转，使其在缺陷存在时失败。我们假设LLM更擅长编写通过测试，而非刻意使其崩溃或失败的测试。实验结果表明，在专门为缺陷可复现测试构建的基准测试集SWT-Bench排行榜上，AssertFlip优于所有已知技术。具体而言，在SWT-Bench-Verified子集上，AssertFlip实现了43.6%的“失败转通过”成功率。

相关内容

大语言模型

关注 65

大语言模型是基于海量文本数据训练的深度学习模型。它不仅能够生成自然语言文本，还能够深入理解文本含义，处理各种自然语言任务，如文本摘要、问答、翻译等。2023年，大语言模型及其在人工智能领域的应用已成为全球科技研究的热点，其在规模上的增长尤为引人注目，参数量已从最初的十几亿跃升到如今的一万亿。参数量的提升使得模型能够更加精细地捕捉人类语言微妙之处，更加深入地理解人类语言的复杂性。在过去的一年里，大语言模型在吸纳新知识、分解复杂任务以及图文对齐等多方面都有显著提升。随着技术的不断成熟，它将不断拓展其应用范围，为人类提供更加智能化和个性化的服务，进一步改善人们的生活和生产方式。

【EMNLP2025】ReCode：基于细粒度检索增强生成的LLM代码修复方法

专知会员服务

10+阅读 · 2025年9月3日

RAG+LLM=？同济大学等最新《大型语言模型的检索增强生成》综述

专知会员服务

111+阅读 · 2023年12月19日

弹药异常检测《使用机器学习进行缺陷表征》最佳论文，MODSIM World 2023

专知会员服务

36+阅读 · 2023年7月22日

生成先验的信号恢复

专知会员服务

22+阅读 · 2023年1月5日