The coexistence of NR-U and Wi-Fi in unlicensed spectrum introduces a system-level resource coordination problem, where heterogeneous channel access mechanisms lead to a significant imbalance in spectrum utilization and degraded Wi-Fi performance. To address this challenge, we propose a policy-driven deep reinforcement learning (DRL) framework for adaptive TXOP control, in which the coexistence process is formulated as a Markov decision process (MDP) and a deep Q-network (DQN) learns control policies through online interaction. A key contribution is the introduction of a policy layer via reward design, enabling explicit control of system-level tradeoffs among fairness, throughput, and quality of service (QoS). Three policies, namely absolute fairness, moderate fairness, and utility-based fairness, are developed to achieve different operating points. Simulation results show that the proposed framework achieves a Jain fairness index above 0.9 under strict fairness control. Compared to absolute fairness, moderate fairness improves aggregate throughput by 68.22%, while the utility-based policy further enhances utility by 177.6%. These results demonstrate that policy-driven control provides a flexible and effective solution for managing tradeoffs in heterogeneous coexistence networks.
翻译:在非授权频段中,NR-U与Wi-Fi的共存引发了系统级资源协调问题:异构信道接入机制导致频谱利用严重失衡及Wi-Fi性能下降。针对这一挑战,我们提出一种策略驱动的深度强化学习(DRL)框架用于自适应TXOP控制,其中共存过程被建模为马尔可夫决策过程(MDP),并通过深度Q网络(DQN)在线学习控制策略。核心贡献在于通过奖励设计引入策略层,使系统能够显式控制公平性、吞吐量和服务质量(QoS)之间的权衡。我们开发了绝对公平、适度公平和基于效用的公平三种策略,以实现不同运行工况。仿真结果表明,在严格公平控制下,所提框架的Jain公平指数可达0.9以上。相较于绝对公平策略,适度公平策略使总吞吐量提升68.22%,而基于效用的策略进一步将效用指标提升177.6%。这些结果证明,策略驱动控制为实现异构共存网络中的权衡管理提供了灵活有效的解决方案。