The complex transmission mechanism of cross-packet hybrid automatic repeat request (XP-HARQ) hinders its optimal system design. To overcome this difficulty, this letter attempts to use the deep reinforcement learning (DRL) to solve the rate selection problem of XP-HARQ over correlated fading channels. In particular, the long term average throughput (LTAT) is maximized by properly choosing the incremental information rate for each HARQ round on the basis of the outdated channel state information (CSI) available at the transmitter. The rate selection problem is first converted into a Markov decision process (MDP), which is then solved by capitalizing on the algorithm of deep deterministic policy gradient (DDPG) with prioritized experience replay. The simulation results finally corroborate the superiority of the proposed XP-HARQ scheme over the conventional HARQ with incremental redundancy (HARQ-IR) and the XP-HARQ with only statistical CSI.
翻译:跨数据包混合自动重传请求(XP-HARQ)的复杂传输机制阻碍了其系统优化设计。为克服这一难题,本文尝试利用深度强化学习(DRL)解决相关衰落信道下XP-HARQ的速率选择问题。具体而言,基于发射端可获取的过时信道状态信息(CSI)合理选择每个HARQ轮次的增量信息速率,以最大化长期平均吞吐量(LTAT)。该速率选择问题首先被转化为马尔可夫决策过程(MDP),随后利用具有优先经验回放的深度确定性策略梯度(DDPG)算法进行求解。仿真结果最终证实了所提出的XP-HARQ方案相较于传统增量冗余HARQ(HARQ-IR)及仅依赖统计CSI的XP-HARQ的优越性。