Cloud computing is the backbone of the digital society. Digital banking, media, communication, gaming, and many others depend on cloud services. Unfortunately, cloud services may fail, leading to damaged services, unhappy users, and perhaps millions of dollars lost for companies. Understanding a cloud service failure requires a detailed report on why and how the service failed. Previous work studies how cloud services fail using logs published by cloud operators. However, information is lacking on how users perceive and experience cloud failures. Therefore, we collect and characterize the data for user-reported cloud failures from Down Detector for three cloud service providers over three years. We count and analyze time patterns in the user reports, and derive failures from those user reports and characterize their duration and interarrival time. We characterize provider-reported cloud failures and compare the results with the characterization of user-reported failures. The comparison reveals the information of how users perceive failures and how much of the failures are reported by cloud service providers. Overall, this work provides a characterization of user- and provider-reported cloud failures and compares them with each other.
翻译:云计算是数字社会的支柱。数字银行、媒体、通信、游戏等诸多领域均依赖于云服务。然而,云服务可能出现故障,导致服务受损、用户不满,甚至使企业蒙受数百万美元的损失。理解云服务故障需要详细报告其发生原因与过程。现有研究通过分析云运营商发布的日志来探究云服务故障,但缺乏关于用户如何感知和体验云故障的信息。为此,我们收集并分析了三年间来自Down Detector平台的三家云服务商的用户报告数据,对用户报告的时间模式进行统计与特征分析,从用户报告中提取故障事件并刻画其持续时长与到达间隔时间。同时,我们对云服务商报告的故障进行特征分析,并将结果与用户报告的故障特征进行对比。对比揭示了用户对故障的感知方式以及云服务商报告的故障覆盖程度。总体而言,本研究提供了用户报告与云服务商报告的云故障特征化分析,并开展了二者的对比研究。