Future wireless networks demand rapid adaptation to highly heterogeneous environments and dynamic task configurations, necessitating a shift from conventional rule-based and optimization-driven radio resource management (RRM) toward artificial intelligence (AI)-driven RRM. AI-driven approaches can learn complex nonlinear relationships, generalize across diverse network conditions and enable real-time, scalable and autonomous decision-making. Among RRM techniques, coordinated multipoint (CoMP) transmission is pivotal for mitigating inter-cell interference and enhancing cell-edge performance, thereby improving quality of experience (QoE) in dense deployments. However, optimal multi-cell selection remains a complex combinatorial challenge as it requires jointly optimizing over many possible serving-cell combinations under dynamic traffic and channel conditions. Despite their success, conventional deep reinforcement learning (DRL) methods such as proximal policy optimization (PPO) suffer from poor sample efficiency, limited generalization, and costly retraining when state and action spaces change. To address these bottlenecks, we propose a Prompt Decision Transformer (PromptDT) based multi-task learning framework capable of learning across diverse network configurations and reformulating multi-cell selection as a sequence modeling problem. By leveraging offline trajectories and task-specific prompts, PromptDT enables scalable learning across diverse network configurations, including varying base stations and user equipment counts, and scheduler policies. Experimental results demonstrate that PromptDT improves QoE by up to 49% in multi-task settings compared to baselines, with performance scaling positively alongside model capacity. Moreover, PromptDT generalizes effectively to unseen tasks, achieving robust few-shot adaptation to new network configurations without retraining or fine-tuning.


翻译:未来无线网络需要快速适应高度异质的环境和动态任务配置,这要求传统基于规则和优化驱动的无线资源管理(RRM)向人工智能(AI)驱动的RRM转变。AI驱动的方法能够学习复杂的非线性关系,在各种网络条件下实现泛化,并支持实时、可扩展和自主的决策。在RRM技术中,协作多点(CoMP)传输对于减轻小区间干扰、提升小区边缘性能至关重要,从而在密集部署中改善用户体验质量(QoE)。然而,最优的多小区选择仍是一个复杂的组合优化问题,因为它需要在动态流量和信道条件下,联合优化众多可能的服务小区组合。尽管传统深度强化学习(DRL)方法(如近端策略优化(PPO))取得了成功,但当状态和动作空间发生变化时,它们仍存在样本效率低、泛化能力有限以及重新训练成本高的问题。为解决这些瓶颈,我们提出了一种基于提示决策Transformer(PromptDT)的多任务学习框架,该框架能够跨不同网络配置进行学习,并将多小区选择问题重新表述为序列建模问题。通过利用离线轨迹和任务特定提示,PromptDT实现了跨不同网络配置(包括变化的基站和用户设备数量以及调度策略)的可扩展学习。实验结果表明,在多任务环境下,与基线方法相比,PromptDT可将QoE提升高达49%,且性能随模型容量的增大而正向扩展。此外,PromptDT能有效泛化到未见过的任务,无需重新训练或微调,即可实现对新型网络配置的鲁棒小样本适应。

0
下载
关闭预览

相关内容

《网络战仿真中的多智能体强化学习》最新42页报告
专知会员服务
48+阅读 · 2023年7月11日
「基于通信的多智能体强化学习」 进展综述
PlaNet 简介:用于强化学习的深度规划网络
谷歌开发者
13+阅读 · 2019年3月16日
Representation Learning on Network 网络表示学习
全球人工智能
10+阅读 · 2017年10月19日
【强化学习】强化学习+深度学习=人工智能
产业智能官
55+阅读 · 2017年8月11日
共享相关任务表征,一文读懂深度神经网络多任务学习
深度学习世界
16+阅读 · 2017年6月23日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
24+阅读 · 2011年12月31日
Arxiv
69+阅读 · 2022年6月13日
VIP会员
最新内容
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
“史诗怒火”行动:现代多域作战的重要节点
专知会员服务
8+阅读 · 7月30日
《下一代无线网络中的多无人机通信资源管理》
相关基金
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
24+阅读 · 2011年12月31日
Top
微信扫码咨询专知VIP会员