In the forthcoming era of 6G networks, characterized by unprecedented data rates, ultra-low latency, and ubiquitous connectivity, effective management of Virtualized Network Functions (VNFs) is essential. VNFs are software-based counterparts of traditional hardware devices that facilitate flexible and scalable service provisioning. Service Function Chains (SFCs), structured as ordered sequences of VNFs, are pivotal in delivering complex network services. Nevertheless, splitting an SFC into multiple segments that are deployed across different network domains or infrastructure locations presents substantial challenges due to the potential heterogeneity of domain characteristic along with quality of service (QoS) constraints and limited visibility of network state. Conventional optimization methods have limited scalability, while existing data-driven approaches struggle to balance efficiency with capturing VNF inter-dependencies in SFCs. To overcome these limitations, we introduce a Transformer-empowered actor-critic framework specifically designed for sequence-aware SFC partitioning. By utilizing the self-attention mechanism, our approach effectively models complex inter-dependencies between VNFs, facilitating coordinated and parallel decision-making processes. Furthermore, to improve training stability and convergence we introduce an $ε$-LoPe exploration strategy as well as Asymptotic Return Normalization. Comprehensive simulation results demonstrate that the proposed methodology outperforms existing state-of-the-art solutions in terms of long-term service acceptance rates, resource utilization, and scalability while achieving fast inference.
翻译:在即将到来的6G网络时代,其特征为空前的数据速率、超低延迟和泛在连接,有效的虚拟化网络功能(VNF)管理至关重要。VNF是传统硬件设备的软件化对应物,支持灵活且可扩展的服务提供。服务功能链(SFC)作为由有序VNF序列构成的结构,在交付复杂网络服务中扮演关键角色。然而,由于域特性的潜在异构性、服务质量(QoS)约束以及网络状态可见性有限,将SFC分割为部署在不同网络域或基础设施位置的多个段带来了重大挑战。传统优化方法的可扩展性有限,而现有数据驱动方法在平衡效率与捕捉SFC中VNF间依赖关系方面面临困难。为克服这些局限,我们引入了一个Transformer赋能的角色-评论框架,专门用于序列感知的SFC分割。通过利用自注意力机制,我们的方法有效地建模了VNF之间的复杂依赖关系,促进了协调的并行决策过程。此外,为提升训练稳定性与收敛性,我们引入了$ε$-LoPe探索策略以及渐近回报归一化(Asymptotic Return Normalization)。综合仿真结果表明,所提出的方法在长期服务接受率、资源利用率和可扩展性方面优于现有最先进解决方案,同时实现快速推理。