The coexistence of NR-U and Wi-Fi in unlicensed spectrum introduces a challenging coexistence management problem, where heterogeneous channel access mechanisms lead to a significant imbalance in spectrum utilization and degraded Wi-Fi performance. To address this challenge, we propose a policy-driven deep reinforcement learning (DRL) framework for adaptive transmission opportunity (TXOP) control, in which the coexistence process is formulated as a Markov decision process (MDP) and a deep Q-network (DQN) learns control policies through online interaction. A key contribution is the introduction of a policy layer via reward design, enabling explicit control of coexistence tradeoffs among fairness, throughput, and utility. Three policies, namely absolute fairness, moderate fairness, and utility-based fairness, are developed to achieve different operating points. Simulation results show that the proposed framework achieves a Jain fairness index above 0.9 under strict fairness control. Compared to absolute fairness, moderate fairness improves aggregate throughput by 68.22%, while the utility-based policy further enhances utility by 177.6%. These results demonstrate that policy-driven control provides a flexible and effective solution for managing tradeoffs in heterogeneous coexistence networks.


翻译:非授权频谱中NR-U与Wi-Fi的共存带来了具有挑战性的共存管理问题:异构信道接入机制将导致频谱利用的显著失衡与Wi-Fi性能下降。为应对这一挑战,我们提出了一种策略驱动的深度强化学习(DRL)框架用于自适应传输机会(TXOP)控制,其中共存过程被建模为马尔可夫决策过程(MDP),并通过深度Q网络(DQN)在在线交互中学习控制策略。核心贡献在于通过奖励设计引入了策略层,从而能够显式控制公平性、吞吐量与效用之间的共存权衡。我们开发了绝对公平、适度公平和基于效用的公平三种策略,以实现不同工作点。仿真结果表明,所提框架在严格公平控制下实现了高于0.9的Jain公平指数。相较于绝对公平,适度公平将聚合吞吐量提升了68.22%,而基于效用的策略则进一步将效用提升了177.6%。这些结果表明,策略驱动的控制为管理异构共存网络中的权衡提供了一种灵活且有效的解决方案。

0
下载
关闭预览

相关内容

网络情报(WI)是网络情报联盟(WIC)的官方期刊,WIC是一个国际组织,致力于促进网络情报时代的合作科研和工业发展。WI寻求与该领域的主要协会和国际会议合作。WI是一份同行评议的期刊,每年出版四期,电子版和纸质版都有。 官网地址:http://dblp.uni-trier.de/db/journals/wias/
《可解释深度强化学习综述》
专知会员服务
40+阅读 · 2025年2月12日
基于强化学习的无人机自组网路由研究综述
专知会员服务
50+阅读 · 2023年9月9日
综述| 当图神经网络遇上强化学习
图与推荐
35+阅读 · 2022年7月1日
当深度强化学习遇见图神经网络
专知
227+阅读 · 2019年10月21日
Self-Attention GAN 中的 self-attention 机制
PaperWeekly
12+阅读 · 2019年3月6日
一文读懂深度适配网络(DAN)
数据派THU
29+阅读 · 2017年7月14日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
VIP会员
最新内容
对抗环境下超视距目标打击的情报支援
专知会员服务
3+阅读 · 今天14:49
《无人机对海面作战影响评估》
专知会员服务
11+阅读 · 7月21日
印度精确打击与指挥架构的断层
专知会员服务
6+阅读 · 7月20日
美空军AI完成F-16战斗机自主空战历史性试飞
专知会员服务
6+阅读 · 7月20日
相关VIP内容
《可解释深度强化学习综述》
专知会员服务
40+阅读 · 2025年2月12日
基于强化学习的无人机自组网路由研究综述
专知会员服务
50+阅读 · 2023年9月9日
相关基金
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
Top
微信扫码咨询专知VIP会员