The placement of Kubernetes control-plane nodes is critical to ensuring cluster reliability, scalability, and performance, and therefore represents a significant deployment challenge in heterogeneous, multi-region environments. Existing initialisation procedures typically select control-plane hosts arbitrarily, without considering node resource capacity or network topology, often leading to suboptimal cluster performance and reduced resilience. Given Kubernetes's status as the de facto standard for container orchestration, there is a need to rigorously evaluate how control-plane node placement influences the overall performance of the cluster operating across multiple regions. This paper advances this goal by introducing an intelligent methodology for selecting control-plane node placement across dynamically selected Cloud-Edge resources spanning multiple regions, as part of an automated orchestration system. More specifically, we propose a reinforcement learning framework based on neural contextual bandits that observes operational performance and learns optimal control-plane placement policies from infrastructure characteristics. Experimental evaluation across several geographically distributed regions and multiple cluster configurations demonstrates substantial performance improvements over several baseline approaches.
翻译:Kubernetes控制平面节点的放置对于确保集群的可靠性、可扩展性和性能至关重要,因此在异构多区域环境中构成重大部署挑战。现有初始化流程通常任意选择控制平面主机,未考虑节点资源容量或网络拓扑,常导致集群性能欠佳和弹性降低。鉴于Kubernetes作为容器编排事实标准(de facto standard)的地位,亟需严谨评估控制平面节点放置如何影响跨多区域集群的整体性能。本文通过提出一种智能方法论,在自动编排系统中为跨多区域的动态选择云边(Cloud-Edge)资源优化控制平面节点放置,从而推进该目标。具体而言,我们提出了基于神经上下文老虎机(neural contextual bandits)的强化学习框架,该框架观测运行性能并从基础设施特征中学习最优控制平面放置策略。跨多个地理分布区域及多种集群配置的实验评估表明,该方法相较于若干基线方法取得了显著性能提升。