We consider the problem of dynamic platoon leader selection, user association, channel assignment, and power allocation on a cellular vehicle-to-everything (C-V2X) based highway, where multiple vehicle-to-vehicle (V2V) and vehicle-to-infrastructure (V2I) links share the frequency resources. There are multiple roadside units (RSUs) on a highway, and vehicles can form platoons, which has been identified as an advanced use case to increase road efficiency. The traditional optimization methods, requiring global channel information at a central controller, are not viable for high-mobility vehicular networks. To deal with this challenge, we propose a distributed multi-agent reinforcement learning (MARL) for resource allocation (RA). Each platoon leader, acting as an agent, can collaborate with other agents for joint sub-band selection and power allocation for its V2V links, and joint user association and power control for its V2I links. Moreover, each platoon can dynamically select the vehicle most suitable to be the platoon leader. We aim to maximize the V2V and V2I packet delivery probability in the desired latency using the deep Q-learning algorithm. Simulation results indicate that our proposed MARL outperforms the centralized hill-climbing algorithm, and platoon leader selection helps to improve both V2V and V2I performance.
翻译:本文研究在基于蜂窝车联网(C-V2X)的高速公路场景中动态编队头车选择、用户关联、信道分配与功率分配问题,其中多条车对车(V2V)与车对基础设施(V2I)链路共享频率资源。高速公路沿线部署多个路侧单元(RSU),车辆可组成编队——这已被确认为提升道路效率的高级应用场景。传统优化方法需要中央控制器获取全局信道信息,难以适用于高移动性车载网络。为应对这一挑战,我们提出一种分布式多智能体强化学习(MARL)方法进行资源分配(RA)。每个编队头车作为智能体,可与其他智能体协作实现V2V链路的联合子带选择与功率分配,以及V2I链路的联合用户关联与功率控制。此外,每个编队可动态选择最适合担任头车的车辆。我们旨在利用深度Q学习算法,在满足预期时延约束下最大化V2V与V2I数据包投递概率。仿真结果表明,本文提出的MARL算法优于集中式爬山算法,且编队头车选择有助于提升V2V与V2I性能。