In recent years, several HPC facilities have started continuous monitoring of their systems and jobs to collect performance-related data for understanding performance and operational efficiency. Such data can be used to optimize the performance of individual jobs and the overall system by creating data-driven models that can predict the performance of jobs waiting in the scheduler queue. In this paper, we model the performance of representative control jobs using longitudinal system-wide monitoring data and machine learning to explore the causes of performance variability. We analyze these prediction models in great detail to identify the features that are dominant predictors of performance. We demonstrate that such models can be application-agnostic and can be used for predicting performance of applications that are not included in training.
翻译:近年来,多个高性能计算设施开始对其系统和作业进行持续监控,以收集与性能相关的数据,从而理解性能和运行效率。通过创建能够预测调度器队列中等待作业性能的数据驱动模型,此类数据可用于优化单个作业及整体系统的性能。在本文中,我们利用纵向全系统监控数据和机器学习对代表性控制作业的性能进行建模,以探究性能变异性的成因。我们深入分析了这些预测模型,以识别对性能具有主导预测作用的特征。我们证明,此类模型可以做到与具体应用无关,并可用于预测训练中未包含的应用的性能。