Software testing is critical for verifying that systems meet specified requirements, yet remains among the most time-consuming and expensive activities in development. Requirements-based test generation allows test cases to be derived early from requirements artifacts, but generating them directly from natural language is challenging due to inherent ambiguity and imprecision. Recent advances in AI, natural language processing (NLP), and large language models (LLMs) have made automating this pipeline increasingly feasible, while introducing new risks including hallucination, reduced traceability, and inconsistent evaluation. This survey addresses four research questions: what AI and NLP techniques have been proposed for generating test cases from natural language requirements; what tools and frameworks support these approaches; how generated test cases are evaluated; and what research gaps remain. Following Kitchenham and Charters' systematic review guidelines, we searched major scholarly databases spanning 2000-2025 and, after applying strict inclusion criteria, identified 21 primary studies. The literature is organized into three evolutionary eras, revealing that no existing approach simultaneously satisfies six key quality dimensions: automation, ambiguity handling, domain applicability, traceability, evaluation thoroughness, and hallucination control. The survey makes three main contributions: a three-era evolutionary synthesis of AI-based test generation; a six-criteria gap analysis showing no current approach fully addresses all quality dimensions; and four actionable research guidelines targeting hallucination, traceability, complexity sensitivity, and compliance.
翻译:软件测试对于验证系统是否满足指定需求至关重要,但仍是开发中最耗时且成本最高的活动之一。基于需求的测试生成允许从需求工件早期推导测试用例,但直接从自然语言生成测试用例因其固有的歧义性和不精确性而充满挑战。近年来人工智能(AI)、自然语言处理(NLP)和大语言模型(LLM)的进展使得自动化这一流程变得愈发可行,但同时也引入了新风险,包括幻觉生成、可追溯性降低以及评估标准不一致。本综述围绕四个研究问题展开:哪些AI与NLP技术已被提出用于从自然语言需求生成测试用例;哪些工具和框架支持这些方法;如何评估所生成的测试用例;以及当前存在哪些研究空白。遵循Kitchenham与Charters的系统性综述指南,我们在覆盖2000-2025年的主要学术数据库中实施检索,经严格纳入标准筛选后,最终确定21篇核心文献。文献被划分为三个进化时代,揭示出现有方法无法同时满足自动化程度、歧义处理、领域适用性、可追溯性、评估全面性与幻觉控制这六个关键质量维度。本综述贡献三点:提出基于AI的测试生成三时代演进综合框架;通过六维度差距分析发现尚无方法能全面满足所有质量指标;针对幻觉生成、可追溯性、复杂度敏感性和合规性提出四项可操作的研究指南。