Reinforcement learning for service orchestration has been the subject of sustained research for over a decade, yet it is not used in production at scale. The usual explanation is that learned controllers degrade under delayed and noisy telemetry, workload shifts, and uncontrolled tenants. We test whether existing evidence supports that explanation. We evaluate three highly influential RL-based orchestration systems spanning resource allocation, DAG scheduling, and autoscaling, using pre-registered predictions about comparative degradation under production-relevant perturbations and paired inference with family-wise error correction. Across the tests, most predicted performance reversals do not occur. Diagnostic analyses show that these outcomes often reflect comparator collapse, artefact limitations, or evaluation choices rather than evidence that learned controllers tolerate the perturbations. One apparent advantage under observation lag is roughly fortyfold compared to a Kubernetes HPA-equivalent controller. Another widely cited result cannot be reconstructed from its released artefact, and the strongest reproducible margin is far smaller than the published results. Conclusions also reverse under changes in perturbation magnitude and evaluation mode. Based on these results and broader patterns in the literature, we identify an institutional problem. Publication and review incentives favour benchmark gains against convenient comparators, even when those gains provide little evidence of deployment performance. We argue that the problem is not solely technical. Rather, it is institutional, so learned orchestration needs production-grade comparators, registered perturbation models, separate operational metrics, and publication criteria that reward reproducible operational evidence. Without these changes, the literature can grow without establishing whether learning improves orchestration.


翻译:针对服务编排的强化学习研究已持续十余年,但至今未在生产环境中大规模应用。普遍解释是学习型控制器在面对延迟与噪声遥测、负载变化及不受控租户时性能会退化。我们通过实验检验该解释是否成立。针对资源分配、DAG调度和自动扩缩容三个领域,我们评估了三个具有高影响力的基于强化学习的编排系统,采用预注册的预测方法比较其在接近生产环境的扰动下的性能退化程度,并使用家族误差校正进行配对统计推断。测试结果显示,大多数预测的性能反转并未发生。诊断分析表明,这些结果往往源于对照组性能崩溃、工具局限性或评估方式选择,而非学习型控制器确实能容忍扰动。在观测延迟扰动下,某个学习型控制器展现出约相当于Kubernetes HPA等效控制器四十倍的优势。另一项被广泛引用的结果无法从其发布的工具中复现,且最大可复现效果幅度远小于原始论文。结论在改变扰动强度与评估模式时也会发生逆转。基于上述结果及文献中的更广泛模式,我们识别出一个制度性问题:发表与评审激励机制更偏向于在便利对照组上取得基准提升,即使这些提升几乎不能证明实际部署性能。我们认为这并非纯技术问题,而是制度性挑战。因此,学习型编排需要面向生产环境的对照组、预注册扰动模型、独立运维指标,以及奖励可复现实验证据的发表标准。若无这些变革,即使文献持续增长,也无法确定强化学习是否真正改善了编排性能。

0
下载
关闭预览

相关内容

面向强化学习的可解释性研究综述
专知会员服务
45+阅读 · 2024年7月30日
持续学习:研究综述
专知会员服务
83+阅读 · 2023年1月30日
【干货书】强化学习Python真实数据与实例应用,110页pdf
专知会员服务
115+阅读 · 2022年10月13日
可解释强化学习,Explainable Reinforcement Learning: A Survey
专知会员服务
133+阅读 · 2020年5月14日
「强化学习可解释性」最新2022综述
专知
12+阅读 · 2022年1月16日
Distributional Soft Actor-Critic (DSAC)强化学习算法的设计与验证
深度强化学习实验室
20+阅读 · 2020年8月11日
最新《多任务学习》综述,39页pdf
专知
28+阅读 · 2020年7月10日
浅谈主动学习(Active Learning)
凡人机器学习
32+阅读 · 2020年6月18日
【强化学习】强化学习/增强学习/再励学习介绍
产业智能官
10+阅读 · 2018年2月23日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
41+阅读 · 2015年12月31日
国家自然科学基金
17+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
31+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2015年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
Arxiv
29+阅读 · 2023年2月10日
VIP会员
最新内容
《边缘计算关键技术分析及美军作战实践应用》
专知会员服务
0+阅读 · 今天14:08
边缘计算的军事应用
专知会员服务
1+阅读 · 今天13:50
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
4+阅读 · 8月8日
《多域冲突比较支持模型》60页
专知会员服务
10+阅读 · 8月7日
相关VIP内容
面向强化学习的可解释性研究综述
专知会员服务
45+阅读 · 2024年7月30日
持续学习:研究综述
专知会员服务
83+阅读 · 2023年1月30日
【干货书】强化学习Python真实数据与实例应用,110页pdf
专知会员服务
115+阅读 · 2022年10月13日
可解释强化学习,Explainable Reinforcement Learning: A Survey
专知会员服务
133+阅读 · 2020年5月14日
相关基金
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
41+阅读 · 2015年12月31日
国家自然科学基金
17+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
31+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2015年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
Top
微信扫码咨询专知VIP会员