In Agile software development, user stories play a vital role in capturing and conveying end-user needs, prioritizing features, and facilitating communication and collaboration within development teams. However, automated methods for evaluating user stories require training in NLP tools and can be time-consuming to develop and integrate. This study explores using ChatGPT for user story quality evaluation and compares its performance with an existing benchmark. Our study shows that ChatGPT's evaluation aligns well with human evaluation, and we propose a ``best of three'' strategy to improve its output stability. We also discuss the concept of trustworthiness in AI and its implications for non-experts using ChatGPT's unprocessed outputs. Our research contributes to understanding the reliability and applicability of AI in user story evaluation and offers recommendations for future research.
翻译:在敏捷软件开发中,用户故事在捕获和传达终端用户需求、确定功能优先级以及促进开发团队内部沟通与协作方面扮演着关键角色。然而,用户故事的自动化评估方法需要对自然语言处理工具进行训练,且开发与集成过程耗时较长。本研究探索了使用ChatGPT进行用户故事质量评估,并将其性能与现有基准进行了比较。研究表明,ChatGPT的评估结果与人类评估高度一致,我们提出了一种"三选最佳"策略以提升其输出稳定性。同时,我们探讨了AI可信度概念及其对非专业人士使用ChatGPT原始输出的启示。本项研究有助于理解AI在用户故事评估中的可靠性与适用性,并为未来研究提供了建议。