Large Language Models (LLMs), such as ChatGPT and BERT, are leading a new AI heatwave due to its human-like conversations with detailed and articulate answers across many domains of knowledge. While LLMs are being quickly applied to many AI application domains, we are interested in the following question: Can safety analysis for safety-critical systems make use of LLMs? To answer, we conduct a case study of Systems Theoretic Process Analysis (STPA) on Automatic Emergency Brake (AEB) systems using ChatGPT. STPA, one of the most prevalent techniques for hazard analysis, is known to have limitations such as high complexity and subjectivity, which this paper aims to explore the use of ChatGPT to address. Specifically, three ways of incorporating ChatGPT into STPA are investigated by considering its interaction with human experts: one-off simplex interaction, recurring simplex interaction, and recurring duplex interaction. Comparative results reveal that: (i) using ChatGPT without human experts' intervention can be inadequate due to reliability and accuracy issues of LLMs; (ii) more interactions between ChatGPT and human experts may yield better results; and (iii) using ChatGPT in STPA with extra care can outperform human safety experts alone, as demonstrated by reusing an existing comparison method with baselines. In addition to making the first attempt to apply LLMs in safety analysis, this paper also identifies key challenges (e.g., trustworthiness concern of LLMs, the need of standardisation) for future research in this direction.
翻译:大语言模型(LLMs),如ChatGPT和BERT,凭借其在多个知识领域提供细致且条理清晰的人类式对话能力,正引领新一轮人工智能热潮。尽管LLMs正快速应用于众多AI应用领域,我们关注以下问题:面向安全关键系统的安全分析能否利用LLMs?为回答此问题,我们以自动紧急制动(AEB)系统为对象,基于ChatGPT开展系统理论过程分析(STPA)案例研究。STPA作为危害分析领域最主流的技术之一,存在复杂度高、主观性强等局限性,本文旨在探索通过ChatGPT缓解这些局限。具体而言,我们研究了三种将ChatGPT融入STPA的方式,并考虑其与人类专家的交互模式:单次单向交互、重复单向交互和重复双向交互。对比结果表明:(i)因LLMs的可靠性及准确性问题,在无人类专家干预下使用ChatGPT存在不足;(ii)ChatGPT与人类专家间增加交互次数可提升效果;(iii)在STPA中使用ChatGPT时,若辅以额外审慎措施,其性能可超越单独的人类安全专家——这一点通过复用现有基准对比方法得到验证。本文不仅首次尝试将LLMs应用于安全分析领域,还指出了该研究方向未来的关键挑战(如LLMs的可信性问题、标准化需求等)。