Recent years have seen the adoption of Machine Learning (ML) techniques to predict the performance of large-scale applications, mostly at a coarse level. In contrast, we propose to use ML techniques for performance prediction at a much finer granularity, namely at the Basic Block (BB) level, which are single entry, single exit code blocks that are used for analysis by the compilers to break down a large code into manageable pieces. We extrapolate the basic block execution counts of GPU applications and use them for predicting the performance for large input sizes from the counts of smaller input sizes. We train a Poisson Neural Network (PNN) model using random input values as well as the lowest input values of the application to learn the relationship between inputs and basic block counts. Experimental results show that the model can accurately predict the basic block execution counts of 16 GPU benchmarks. We achieve an accuracy of 93.5% in extrapolating the basic block counts for large input sets when trained on smaller input sets and an accuracy of 97.7% in predicting basic block counts on random instances. In a case study, we apply the ML model to CUDA GPU benchmarks for performance prediction across a spectrum of applications. We use a variety of metrics for evaluation, including global memory requests and the active cycles of tensor cores, ALU, and FMA units. Results demonstrate the model's capability of predicting the performance of large datasets with an average error rate of 0.85% and 0.17% for global and shared memory requests, respectively. Additionally, to address the utilization of the main functional units in Ampere architecture GPUs, we calculate the active cycles for tensor cores, ALU, FMA, and FP64 units and achieve an average error of 2.3% and 10.66% for ALU and FMA units while the maximum observed error across all tested applications and units reaches 18.5%.
翻译:近年来,机器学习(ML)技术被广泛应用于大规模应用性能预测,但主要集中在粗粒度层面。相比之下,我们提出在更细粒度层面——即基本块(BB)层面(编译器为将大型代码分解为可管理片段而使用的单入口单出口代码块)应用机器学习技术进行性能预测。我们通过外推GPU应用的基本块执行计数,从较小输入规模的计数预测大规模输入的性能。采用随机输入值与应用最小输入值训练泊松神经网络(PNN)模型,以学习输入与基本块计数之间的关联。实验结果表明,该模型能准确预测16个GPU基准测试的基本块执行计数:在小输入集训练下,对大规模输入集的基本块计数外推准确率达93.5%;在随机实例预测中,基本块计数准确率达97.7%。通过案例研究,我们将该机器学习模型应用于CUDA GPU基准测试,实现跨应用谱系的性能预测。采用包括全局内存请求、张量核心、ALU及FMA单元活跃周期在内的多维度评估指标。结果显示,模型对全局与共享内存请求的大数据集性能预测平均误差率分别为0.85%和0.17%。此外,针对Ampere架构GPU主要功能单元利用率,我们计算了张量核心、ALU、FMA及FP64单元的活跃周期,其中ALU与FMA单元的平均误差分别为2.3%和10.66%,所有测试应用与单元的最大观测误差为18.5%。