Reinforcement learning (RL) has emerged as a promising paradigm for Internet congestion control, achieving higher link utilization than classical heuristics. However, RL-based controllers trained in single-flow environments are not guaranteed to share bandwidth equitably when deployed in multi-flow networks. This paper investigates the fairness properties of Aurora~\cite{jay2019aurora}, a state-of-the-art deep RL congestion controller, and evaluates three post-hoc fairness strategies that preserve Aurora's RL architecture: \emph{reward shaping} (Strategy~A), \emph{observation augmentation} (Strategy~B), and \emph{loss-sensitivity tuning} (Strategy~C). Using a custom shared-bottleneck simulator and Jain's fairness index as the primary metric, we find that modest reward shaping achieves the best fairness while preserving aggregate throughput. All strategies maintain the total bandwidth budget with fairness being achieved through redistribution, not reduction. Beyond the 2-flow homogeneous setting, an extended evaluation across mixed Aurora--CUBIC competition and dynamic flow entry/exit scenarios shows that Strategy~C's loss-sensitivity emerges as the most TCP-friendly mechanism, while Strategy~B is the most stable through dynamic flow-set changes.
翻译:强化学习已成为互联网拥塞控制的一种有前景的范式,相比经典启发式方法,它能实现更高的链路利用率。然而,在单流环境中训练的基于强化学习的控制器,当部署到多流网络时,无法保证公平地共享带宽。本文研究了最先进的深度强化学习拥塞控制器Aurora~\cite{jay2019aurora}的公平性特性,并评估了三种保留Aurora强化学习架构的事后公平性策略:\emph{奖励塑造}(策略A)、\emph{观测增强}(策略B)和\emph{损失敏感度调整}(策略C)。通过使用自定义的共享瓶颈模拟器并以Jain公平性指数作为主要指标,我们发现在保持总吞吐量的前提下,适度的奖励塑造能达到最佳的公平性。所有策略都维持了总带宽预算,公平性是通过重新分配而非减少带宽实现的。在2流同质环境之外,针对混合Aurora-CUBIC竞争和动态流进出场景的扩展评估表明,策略C的损失敏感度机制对TCP最友好,而策略B在动态流集合变化时最为稳定。