Simultaneously transmitting and reflecting reconfigurable intelligent surfaces (STAR-RISs) is a promising passive device that contributes to a full-space coverage via transmitting and reflecting the incident signal simultaneously. As a new paradigm in wireless communications, how to analyze the coverage and capacity performance of STAR-RISs becomes essential but challenging. To solve the coverage and capacity optimization (CCO) problem in STAR-RIS assisted networks, a multi-objective proximal policy optimization (MO-PPO) algorithm is proposed to handle long-term benefits than conventional optimization algorithms. To strike a balance between each objective, the MO-PPO algorithm provides a set of optimal solutions to form a Pareto front (PF), where any solution on the PF is regarded as an optimal result. Moreover, in order to improve the performance of the MO-PPO algorithm, two update strategies, i.e., action-value-based update strategy (AVUS) and loss function-based update strategy (LFUS), are investigated. For the AVUS, the improved point is to integrate the action values of both coverage and capacity and then update the loss function. For the LFUS, the improved point is only to assign dynamic weights for both loss functions of coverage and capacity, while the weights are calculated by a min-norm solver at every update. The numerical results demonstrated that the investigated update strategies outperform the fixed weights MO optimization algorithms in different cases, which includes a different number of sample grids, the number of STAR-RISs, the number of elements in the STAR-RISs, and the size of STAR-RISs. Additionally, the STAR-RIS assisted networks achieve better performance than conventional wireless networks without STAR-RISs. Moreover, with the same bandwidth, millimeter wave is able to provide higher capacity than sub-6 GHz, but at a cost of smaller coverage.
翻译:同时透射和反射可重构智能表面(STAR-RISs)是一种有前途的无源器件,通过同时透射和反射入射信号实现全空间覆盖。作为无线通信中的新范式,如何分析STAR-RISs的覆盖和容量性能变得至关重要但充满挑战。为解决STAR-RIS辅助网络中的覆盖与容量优化(CCO)问题,提出了一种多目标近端策略优化(MO-PPO)算法,以处理传统优化算法难以应对的长期收益问题。为平衡各目标,MO-PPO算法提供一组最优解形成帕累托前沿(PF),其中PF上的任一解均视为最优结果。此外,为提升MO-PPO算法性能,研究了两种更新策略,即基于动作值的更新策略(AVUS)和基于损失函数的更新策略(LFUS)。AVUS的改进点在于整合覆盖与容量的动作值后更新损失函数;LFUS的改进点则仅对覆盖和容量的损失函数分配动态权重,每次更新时通过最小范数求解器计算权重。数值结果表明,所研究的更新策略在不同场景(包括不同采样网格数、STAR-RIS数量、STAR-RIS单元数及STAR-RIS尺寸)下均优于固定权重的多目标优化算法。此外,STAR-RIS辅助网络相较于无STAR-RIS的传统无线网络实现了更优性能。同时,在相同带宽下,毫米波能够提供比sub-6 GHz更高的容量,但以牺牲覆盖范围为代价。