In recent years, the transition to cloud-based platforms in the IT sector has emphasized the significance of cloud incident root cause analysis to ensure service reliability and maintain customer trust. Central to this process is the efficient determination of root causes, a task made challenging due to the complex nature of contemporary cloud infrastructures. Despite the proliferation of AI-driven tools for root cause identification, their applicability remains limited by the inconsistent quality of their outputs. This paper introduces a method for enhancing confidence estimation in root cause analysis tools by prompting retrieval-augmented large language models (LLMs). This approach operates in two phases. Initially, the model evaluates its confidence based on historical incident data, considering its assessment of the evidence strength. Subsequently, the model reviews the root cause generated by the predictor. An optimization step then combines these evaluations to determine the final confidence assignment. Experimental results illustrate that our method enables the model to articulate its confidence effectively, providing a more calibrated score. We address research questions evaluating the ability of our method to produce calibrated confidence scores using LLMs, the impact of domain-specific retrieved examples on confidence estimates, and its potential generalizability across various root cause analysis models. Through this, we aim to bridge the confidence estimation gap, aiding on-call engineers in decision-making and bolstering the efficiency of cloud incident management.
翻译:近年来,IT行业向云平台迁移的趋势凸显了云事故根因分析对保障服务可靠性和维护客户信任的重要性。该过程的核心在于高效确定根因,而现代云基础设施的复杂性使这一任务充满挑战。尽管基于人工智能的根因识别工具日益普及,但其输出质量的不一致性仍制约着实际应用。本文提出一种通过提示增强型检索大语言模型来提升根因分析工具置信度估计的方法。该方法分两阶段运作:首先,模型基于历史事故数据评估自身置信度,并考虑其对证据强度的判断;随后,模型审查预测器产生的根因,并通过优化步骤综合这些评估以确定最终置信度赋值。实验结果表明,我们的方法能使模型有效表达其置信度,生成更校准的评分。我们针对以下研究问题展开探讨:该方法能否利用大语言模型生成校准的置信度分数、领域特定检索示例对置信度估计的影响、以及其在不同根因分析模型间的潜在泛化能力。通过本研究,我们旨在弥合置信度估计鸿沟,协助值班工程师进行决策,并提升云事故管理效率。