Massive random access of devices in the emerging Open Radio Access Network (O-RAN) brings great challenge to the access control and management. Exploiting the bursting nature of the access requests, sparse active user detection (SAUD) is an efficient enabler towards efficient access management, but the sparsity might be deteriorated in case of uncoordinated massive access requests. To dynamically preserve the sparsity of access requests, a reinforcement-learning (RL)-assisted scheme of closed-loop access control utilizing the access class barring technique is proposed, where the RL policy is determined through continuous interaction between the RL agent, i.e., a next generation node base (gNB), and the environment. The proposed scheme can be implemented by the near-real-time RAN intelligent controller (near-RT RIC) in O-RAN, supporting rapid switching between heterogeneous vertical applications, such as mMTC and uRLLC services. Moreover, a data-driven scheme of deep-RL-assisted SAUD is proposed to resolve highly complex environments with continuous and high-dimensional state and action spaces, where a replay buffer is applied for automatic large-scale data collection. An actor-critic framework is formulated to incorporate the strategy-learning modules into the near-RT RIC. Simulation results show that the proposed schemes can achieve superior performance in both access efficiency and user detection accuracy over the benchmark scheme for different heterogeneous services with massive access requests.
翻译:在新兴的开放无线接入网(O-RAN)中,设备的大规模随机接入对接入控制与管理提出了巨大挑战。利用接入请求的突发特性,稀疏活跃用户检测(SAUD)是实现高效接入管理的有效手段,但未协调的大规模接入请求可能恶化稀疏性。为动态保持接入请求的稀疏性,提出了一种基于强化学习(RL)的闭环接入控制方案,该方案利用接入等级限制技术,通过RL智能体(即下一代基站gNB)与环境的持续交互来确定RL策略。所提方案可由O-RAN中的近实时RAN智能控制器(near-RT RIC)实现,支持异构垂直应用(如mMTC和uRLLC业务)间的快速切换。此外,为解决连续高维状态与动作空间的复杂环境,提出了一种数据驱动的深度强化学习辅助SAUD方案,其中采用经验回放缓冲区实现大规模自动数据收集。构建了演员-评论家框架,将策略学习模块集成至near-RT RIC中。仿真结果表明,在处理具有大规模接入请求的不同异构业务时,所提方案在接入效率与用户检测精度上均优于基准方案。