Reinforcement learning (RL) has emerged as a promising paradigm for Internet congestion control, achieving higher link utilization than classical heuristics. However, RL-based controllers trained in single-flow environments are not guaranteed to share bandwidth equitably when deployed in multi-flow networks. This paper investigates the fairness properties of Aurora~\cite{jay2019aurora}, a state-of-the-art deep RL congestion controller, and evaluates three post-hoc fairness strategies that preserve Aurora's RL architecture: \emph{reward shaping} (Strategy~A), \emph{observation augmentation} (Strategy~B), and \emph{loss-sensitivity tuning} (Strategy~C). Using a custom shared-bottleneck simulator and Jain's fairness index as the primary metric, we find that modest reward shaping achieves the best fairness while preserving aggregate throughput. All strategies maintain the total bandwidth budget with fairness being achieved through redistribution, not reduction. Beyond the 2-flow homogeneous setting, an extended evaluation across mixed Aurora--CUBIC competition and dynamic flow entry/exit scenarios shows that Strategy~C's loss-sensitivity emerges as the most TCP-friendly mechanism, while Strategy~B is the most stable through dynamic flow-set changes.


翻译:强化学习已成为互联网拥塞控制的一种有前景的范式,相比经典启发式方法,它能实现更高的链路利用率。然而,在单流环境中训练的基于强化学习的控制器,当部署到多流网络时,无法保证公平地共享带宽。本文研究了最先进的深度强化学习拥塞控制器Aurora~\cite{jay2019aurora}的公平性特性,并评估了三种保留Aurora强化学习架构的事后公平性策略:\emph{奖励塑造}(策略A)、\emph{观测增强}(策略B)和\emph{损失敏感度调整}(策略C)。通过使用自定义的共享瓶颈模拟器并以Jain公平性指数作为主要指标,我们发现在保持总吞吐量的前提下,适度的奖励塑造能达到最佳的公平性。所有策略都维持了总带宽预算,公平性是通过重新分配而非减少带宽实现的。在2流同质环境之外,针对混合Aurora-CUBIC竞争和动态流进出场景的扩展评估表明,策略C的损失敏感度机制对TCP最友好,而策略B在动态流集合变化时最为稳定。

0
下载
关闭预览

相关内容

【NTU博士论文】基于协作式多智能体强化学习的决策制定
多智能体强化学习控制与决策研究综述
专知会员服务
50+阅读 · 2024年11月23日
自动驾驶中的多智能体强化学习综述
专知会员服务
48+阅读 · 2024年8月20日
面向强化学习的可解释性研究综述
专知会员服务
45+阅读 · 2024年7月30日
基于学习机制的多智能体强化学习综述
专知会员服务
64+阅读 · 2024年4月16日
基于通信的多智能体强化学习进展综述
专知会员服务
113+阅读 · 2022年11月12日
「基于通信的多智能体强化学习」 进展综述
基于模型的强化学习综述
专知
42+阅读 · 2022年7月13日
「强化学习可解释性」最新2022综述
专知
12+阅读 · 2022年1月16日
【综述】多智能体强化学习算法理论研究
深度强化学习实验室
16+阅读 · 2020年9月9日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
24+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
12+阅读 · 2008年12月31日
VIP会员
最新内容
俄乌无人机战争的六大启示
专知会员服务
2+阅读 · 今天7:07
《无人机空中监控:通信实验洞察》
专知会员服务
1+阅读 · 今天7:05
从采集到决策:美军视角下的战术情报范式重构
《履带式无人地面战车技术发展现状》
专知会员服务
5+阅读 · 8月2日
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
5+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
13+阅读 · 7月31日
相关基金
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
24+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
12+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员