Deploying adversarially robust machine learning systems requires continuous trade-offs between robustness, cost, and latency. We present an autonomic decision-support framework providing a quantitative foundation for adaptive hardware selection and hyper-parameter tuning in cloud-native deep learning. The framework applies accelerated failure time (AFT) models to quantify the effect of hardware choice, batch size, epochs, and validation accuracy on model survival time. This framework can be naturally integrated into an autonomic control loop (monitor--analyse--plan--execute, MAPE-K), where system metrics such as cost, robustness, and latency are continuously evaluated and used to adapt model configurations and hardware selection. Experiments across three GPU architectures confirm the framework is both sound and cost-effective: the Nvidia L4 yields a 20% increase in adversarial survival time while costing 75% less than the V100, demonstrating that expensive hardware does not necessarily improve robustness. The analysis further reveals that model inference latency is a stronger predictor of adversarial robustness than training time or hardware configuration.
翻译:部署具有对抗鲁棒性的机器学习系统需要在鲁棒性、成本和延迟之间持续权衡。我们提出了一种自主决策支持框架,为云原生深度学习中的自适应硬件选择和超参数调优提供定量基础。该框架应用加速失效时间(AFT)模型,量化硬件选择、批大小、训练轮次和验证准确率对模型生存时间的影响。该框架可自然集成到自主控制回路(监控-分析-规划-执行,MAPE-K)中,持续评估成本、鲁棒性和延迟等系统指标,并据此调整模型配置与硬件选择。在三种GPU架构上的实验证实该框架兼具合理性与成本效益:Nvidia L4在对抗生存时间上提升20%,同时成本比V100降低75%,表明昂贵硬件未必能提升鲁棒性。分析进一步揭示,模型推理延迟对对抗鲁棒性的预测能力优于训练时间或硬件配置。