Reinforcement learning for service orchestration has been the subject of sustained research for over a decade, yet it is not used in production at scale. The usual explanation is that learned controllers degrade under delayed and noisy telemetry, workload shifts, and uncontrolled tenants. We test whether existing evidence supports that explanation. We evaluate three highly influential RL-based orchestration systems spanning resource allocation, DAG scheduling, and autoscaling, using pre-registered predictions about comparative degradation under production-relevant perturbations and paired inference with family-wise error correction. Across the tests, most predicted performance reversals do not occur. Diagnostic analyses show that these outcomes often reflect comparator collapse, artefact limitations, or evaluation choices rather than evidence that learned controllers tolerate the perturbations. One apparent advantage under observation lag is roughly fortyfold compared to a Kubernetes HPA-equivalent controller. Another widely cited result cannot be reconstructed from its released artefact, and the strongest reproducible margin is far smaller than the published results. Conclusions also reverse under changes in perturbation magnitude and evaluation mode. Based on these results and broader patterns in the literature, we identify an institutional problem. Publication and review incentives favour benchmark gains against convenient comparators, even when those gains provide little evidence of deployment performance. We argue that the problem is not solely technical. Rather, it is institutional, so learned orchestration needs production-grade comparators, registered perturbation models, separate operational metrics, and publication criteria that reward reproducible operational evidence. Without these changes, the literature can grow without establishing whether learning improves orchestration.
翻译:针对服务编排的强化学习研究已持续十余年,但至今未在生产环境中大规模应用。普遍解释是学习型控制器在面对延迟与噪声遥测、负载变化及不受控租户时性能会退化。我们通过实验检验该解释是否成立。针对资源分配、DAG调度和自动扩缩容三个领域,我们评估了三个具有高影响力的基于强化学习的编排系统,采用预注册的预测方法比较其在接近生产环境的扰动下的性能退化程度,并使用家族误差校正进行配对统计推断。测试结果显示,大多数预测的性能反转并未发生。诊断分析表明,这些结果往往源于对照组性能崩溃、工具局限性或评估方式选择,而非学习型控制器确实能容忍扰动。在观测延迟扰动下,某个学习型控制器展现出约相当于Kubernetes HPA等效控制器四十倍的优势。另一项被广泛引用的结果无法从其发布的工具中复现,且最大可复现效果幅度远小于原始论文。结论在改变扰动强度与评估模式时也会发生逆转。基于上述结果及文献中的更广泛模式,我们识别出一个制度性问题:发表与评审激励机制更偏向于在便利对照组上取得基准提升,即使这些提升几乎不能证明实际部署性能。我们认为这并非纯技术问题,而是制度性挑战。因此,学习型编排需要面向生产环境的对照组、预注册扰动模型、独立运维指标,以及奖励可复现实验证据的发表标准。若无这些变革,即使文献持续增长,也无法确定强化学习是否真正改善了编排性能。