Context: Requirements engineering researchers have been experimenting with machine learning and deep learning approaches for a range of RE tasks, such as requirements classification, requirements tracing, ambiguity detection, and modelling. However, most of today's ML/DL approaches are based on supervised learning techniques, meaning that they need to be trained using a large amount of task-specific labelled training data. This constraint poses an enormous challenge to RE researchers, as the lack of labelled data makes it difficult for them to fully exploit the benefit of advanced ML/DL technologies. Objective: This paper addresses this problem by showing how a zero-shot learning approach can be used for requirements classification without using any labelled training data. We focus on the classification task because many RE tasks can be framed as classification problems. Method: The ZSL approach used in our study employs contextual word-embeddings and transformer-based language models. We demonstrate this approach through a series of experiments to perform three classification tasks: (1)FR/NFR: classification functional requirements vs non-functional requirements; (2)NFR: identification of NFR classes; (3)Security: classification of security vs non-security requirements. Results: The study shows that the ZSL approach achieves an F1 score of 0.66 for the FR/NFR task. For the NFR task, the approach yields F1~0.72-0.80, considering the most frequent classes. For the Security task, F1~0.66. All of the aforementioned F1 scores are achieved with zero-training efforts. Conclusion: This study demonstrates the potential of ZSL for requirements classification. An important implication is that it is possible to have very little or no training data to perform classification tasks. The proposed approach thus contributes to the solution of the long-standing problem of data shortage in RE.
翻译:背景:需求工程研究人员一直在尝试将机器学习与深度学习方法应用于一系列RE任务,如需求分类、需求追踪、歧义检测及建模。然而,当前多数ML/DL方法基于监督学习技术,这意味着它们需要利用大量任务特定的标记训练数据进行训练。这一限制给RE研究人员带来了巨大挑战,因为标记数据的缺乏使他们难以充分利用先进ML/DL技术的优势。目标:本文通过展示如何在不使用任何标记训练数据的情况下,利用零样本学习方法进行需求分类来解决这一问题。我们聚焦于分类任务,因为许多RE任务可被构建为分类问题。方法:本研究中采用的ZSL方法使用了上下文词嵌入和基于Transformer的语言模型。我们通过一系列实验演示了该方法在三个分类任务中的应用:(1)FR/NFR:功能需求与非功能需求分类;(2)NFR:NFR类别识别;(3)Security:安全需求与非安全需求分类。结果:研究表明,ZSL方法在FR/NFR任务上取得了0.66的F1分数。在NFR任务中,考虑最频繁类别时,该方法获得的F1分数约为0.72-0.80。在Security任务中,F1分数约为0.66。所有上述F1分数均在零训练投入下实现。结论:本研究展示了ZSL在需求分类中的潜力。一个重要启示是,可能仅需极少甚至无需训练数据即可执行分类任务。因此,所提出的方法有助于解决RE领域长期存在的数据短缺问题。