Future wireless networks demand rapid adaptation to highly heterogeneous environments and dynamic task configurations, necessitating a shift from conventional rule-based and optimization-driven radio resource management (RRM) toward artificial intelligence (AI)-driven RRM. AI-driven approaches can learn complex nonlinear relationships, generalize across diverse network conditions and enable real-time, scalable and autonomous decision-making. Among RRM techniques, coordinated multipoint (CoMP) transmission is pivotal for mitigating inter-cell interference and enhancing cell-edge performance, thereby improving quality of experience (QoE) in dense deployments. However, optimal multi-cell selection remains a complex combinatorial challenge as it requires jointly optimizing over many possible serving-cell combinations under dynamic traffic and channel conditions. Despite their success, conventional deep reinforcement learning (DRL) methods such as proximal policy optimization (PPO) suffer from poor sample efficiency, limited generalization, and costly retraining when state and action spaces change. To address these bottlenecks, we propose a Prompt Decision Transformer (PromptDT) based multi-task learning framework capable of learning across diverse network configurations and reformulating multi-cell selection as a sequence modeling problem. By leveraging offline trajectories and task-specific prompts, PromptDT enables scalable learning across diverse network configurations, including varying base stations and user equipment counts, and scheduler policies. Experimental results demonstrate that PromptDT improves QoE by up to 49% in multi-task settings compared to baselines, with performance scaling positively alongside model capacity. Moreover, PromptDT generalizes effectively to unseen tasks, achieving robust few-shot adaptation to new network configurations without retraining or fine-tuning.
翻译:未来无线网络需要快速适应高度异质的环境和动态任务配置,这要求传统基于规则和优化驱动的无线资源管理(RRM)向人工智能(AI)驱动的RRM转变。AI驱动的方法能够学习复杂的非线性关系,在各种网络条件下实现泛化,并支持实时、可扩展和自主的决策。在RRM技术中,协作多点(CoMP)传输对于减轻小区间干扰、提升小区边缘性能至关重要,从而在密集部署中改善用户体验质量(QoE)。然而,最优的多小区选择仍是一个复杂的组合优化问题,因为它需要在动态流量和信道条件下,联合优化众多可能的服务小区组合。尽管传统深度强化学习(DRL)方法(如近端策略优化(PPO))取得了成功,但当状态和动作空间发生变化时,它们仍存在样本效率低、泛化能力有限以及重新训练成本高的问题。为解决这些瓶颈,我们提出了一种基于提示决策Transformer(PromptDT)的多任务学习框架,该框架能够跨不同网络配置进行学习,并将多小区选择问题重新表述为序列建模问题。通过利用离线轨迹和任务特定提示,PromptDT实现了跨不同网络配置(包括变化的基站和用户设备数量以及调度策略)的可扩展学习。实验结果表明,在多任务环境下,与基线方法相比,PromptDT可将QoE提升高达49%,且性能随模型容量的增大而正向扩展。此外,PromptDT能有效泛化到未见过的任务,无需重新训练或微调,即可实现对新型网络配置的鲁棒小样本适应。