Context: Machine learning (ML) may enable effective automated test generation. Objective: We characterize emerging research, examining testing practices, researcher goals, ML techniques applied, evaluation, and challenges. Methods: We perform a systematic mapping on a sample of 124 publications. Results: ML generates input for system, GUI, unit, performance, and combinatorial testing or improves the performance of existing generation methods. ML is also used to generate test verdicts, property-based, and expected output oracles. Supervised learning - often based on neural networks - and reinforcement learning - often based on Q-learning - are common, and some publications also employ unsupervised or semi-supervised learning. (Semi-/Un-)Supervised approaches are evaluated using both traditional testing metrics and ML-related metrics (e.g., accuracy), while reinforcement learning is often evaluated using testing metrics tied to the reward function. Conclusion: Work-to-date shows great promise, but there are open challenges regarding training data, retraining, scalability, evaluation complexity, ML algorithms employed - and how they are applied - benchmarks, and replicability. Our findings can serve as a roadmap and inspiration for researchers in this field.
翻译:语境:机器学习(ML)可能实现有效的自动化测试生成。目标:我们梳理新兴研究,考察测试实践、研究者目标、所用ML技术、评估方法及挑战。方法:对124篇出版物样本进行系统映射分析。结果:ML用于生成系统测试、GUI测试、单元测试、性能测试和组合测试的输入,或改进现有生成方法的性能。ML还被用于生成测试判决、基于属性的预言以及预期输出预言。监督学习(常基于神经网络)和强化学习(常基于Q-学习)较为常见,部分研究也采用无监督或半监督学习。(半/无)监督方法使用传统测试指标和ML相关指标(如准确率)进行评估,而强化学习通常采用与奖励函数相关的测试指标。结论:现有工作展现出巨大潜力,但在训练数据、重新训练、可扩展性、评估复杂度、所用ML算法及其应用方式、基准测试和可复现性方面仍存在开放挑战。我们的发现可为该领域研究者提供路线图与启示。