Due to the growing complexity of modern data centers, failures are not uncommon any more. Therefore, fault tolerance mechanisms play a vital role in fulfilling the availability requirements. Multiple availability models have been proposed to assess compute systems, among which Bayesian network models have gained popularity in industry and research due to its powerful modeling formalism. In particular, this work focuses on assessing the availability of redundant and replicated cloud computing services with Bayesian networks. So far, research on availability has only focused on modeling either infrastructure or communication failures in Bayesian networks, but have not considered both simultaneously. This work addresses practical modeling challenges of assessing the availability of large-scale redundant and replicated services with Bayesian networks, including cascading and common-cause failures from the surrounding infrastructure and communication network. In order to ease the modeling task, this paper introduces a high-level modeling formalism to build such a Bayesian network automatically. Performance evaluations demonstrate the feasibility of the presented Bayesian network approach to assess the availability of large-scale redundant and replicated services. This model is not only applicable in the domain of cloud computing it can also be applied for general cases of local and geo-distributed systems.
翻译:由于现代数据中心的日益复杂性,故障已不再罕见。因此,容错机制在满足可用性要求方面发挥着关键作用。为评估计算系统已提出多种可用性模型,其中贝叶斯网络模型因其强大的建模形式在工业界和学术界广受欢迎。本研究特别聚焦于利用贝叶斯网络评估冗余与复制云服务的可用性。迄今为止,针对可用性的研究仅侧重于在贝叶斯网络中建模基础设施或通信故障,尚未同时考虑两者。本文探讨了利用贝叶斯网络评估大规模冗余与复制服务可用性的实际建模挑战,包括来自周边基础设施及通信网络的级联故障与共因故障。为简化建模任务,本文引入了一种高阶建模形式以自动构建此类贝叶斯网络。性能评估证明了所提出的贝叶斯网络方法在评估大规模冗余与复制服务可用性方面的可行性。该模型不仅适用于云计算领域,也可应用于本地及地理分布式系统的一般场景。