Artificial Intelligence for IT operations (AIOps) aims to combine the power of AI with the big data generated by IT Operations processes, particularly in cloud infrastructures, to provide actionable insights with the primary goal of maximizing availability. There are a wide variety of problems to address, and multiple use-cases, where AI capabilities can be leveraged to enhance operational efficiency. Here we provide a review of the AIOps vision, trends challenges and opportunities, specifically focusing on the underlying AI techniques. We discuss in depth the key types of data emitted by IT Operations activities, the scale and challenges in analyzing them, and where they can be helpful. We categorize the key AIOps tasks as - incident detection, failure prediction, root cause analysis and automated actions. We discuss the problem formulation for each task, and then present a taxonomy of techniques to solve these problems. We also identify relatively under explored topics, especially those that could significantly benefit from advances in AI literature. We also provide insights into the trends in this field, and what are the key investment opportunities.
翻译:人工智能运维(AIOps)旨在将AI能力与IT运维流程(尤其在云基础设施中)产生的大数据相结合,提供可操作洞察,核心目标是最大化系统可用性。当前存在多样化的待解决问题及多种应用场景,可借助AI技术提升运维效率。本文对AIOps的愿景、趋势、挑战与机遇进行了综述,重点聚焦底层AI技术。我们深入剖析了IT运维活动产生的主要数据类型、分析这些数据所面临的规模与挑战,以及其潜在应用价值。将关键AIOps任务归纳为:事件检测、故障预测、根因分析与自动化处置。针对每项任务阐述问题建模方法,并构建解决这些问题的技术分类体系。同时识别了相对未被充分探索的研究方向,特别是那些可能受益于AI领域最新进展的课题。最后,我们对该领域的发展趋势及核心投资机遇进行了展望。