The coexistence of NR-U and Wi-Fi in unlicensed spectrum introduces a challenging coexistence management problem, where heterogeneous channel access mechanisms lead to a significant imbalance in spectrum utilization and degraded Wi-Fi performance. To address this challenge, we propose a policy-driven deep reinforcement learning (DRL) framework for adaptive transmission opportunity (TXOP) control, in which the coexistence process is formulated as a Markov decision process (MDP) and a deep Q-network (DQN) learns control policies through online interaction. A key contribution is the introduction of a policy layer via reward design, enabling explicit control of coexistence tradeoffs among fairness, throughput, and utility. Three policies, namely absolute fairness, moderate fairness, and utility-based fairness, are developed to achieve different operating points. Simulation results show that the proposed framework achieves a Jain fairness index above 0.9 under strict fairness control. Compared to absolute fairness, moderate fairness improves aggregate throughput by 68.22%, while the utility-based policy further enhances utility by 177.6%. These results demonstrate that policy-driven control provides a flexible and effective solution for managing tradeoffs in heterogeneous coexistence networks.
翻译:非授权频谱中NR-U与Wi-Fi的共存带来了具有挑战性的共存管理问题:异构信道接入机制将导致频谱利用的显著失衡与Wi-Fi性能下降。为应对这一挑战,我们提出了一种策略驱动的深度强化学习(DRL)框架用于自适应传输机会(TXOP)控制,其中共存过程被建模为马尔可夫决策过程(MDP),并通过深度Q网络(DQN)在在线交互中学习控制策略。核心贡献在于通过奖励设计引入了策略层,从而能够显式控制公平性、吞吐量与效用之间的共存权衡。我们开发了绝对公平、适度公平和基于效用的公平三种策略,以实现不同工作点。仿真结果表明,所提框架在严格公平控制下实现了高于0.9的Jain公平指数。相较于绝对公平,适度公平将聚合吞吐量提升了68.22%,而基于效用的策略则进一步将效用提升了177.6%。这些结果表明,策略驱动的控制为管理异构共存网络中的权衡提供了一种灵活且有效的解决方案。