Performance tuning, software/hardware co-design, and job scheduling are among the many tasks that rely on models to predict application performance. We propose and evaluate low rank tensor decomposition for modeling application performance. We discretize the input and configuration domain of an application using regular grids. Application execution times mapped within grid-cells are averaged and represented by tensor elements. We show that low-rank canonical-polyadic (CP) tensor decomposition is effective in approximating these tensors. We further show that this decomposition enables accurate extrapolation of unobserved regions of an application's parameter space. We then employ tensor completion to optimize a CP decomposition given a sparse set of observed runtimes. We consider alternative piecewise/grid-based models and supervised learning models for six applications and demonstrate that CP decomposition optimized using tensor completion offers higher prediction accuracy and memory-efficiency for high-dimensional applications.
翻译:性能调优、软硬件协同设计及任务调度等众多任务依赖于对应用程序性能进行预测的模型。我们提出并评估了基于低秩张量分解的应用程序性能建模方法。我们使用规则网格对应用的输入和配置域进行离散化处理。映射到网格单元内的应用执行时间经过平均后以张量元素表示。研究表明,低秩典型多向(CP)张量分解能够有效逼近这些张量。进一步的分析显示,该分解方法可实现对应用参数空间中未观测区域的精确外推。随后,我们采用张量补全技术,基于稀疏的观测运行时数据优化CP分解。我们针对六个应用比较了分段/网格模型及监督学习模型,实验证明:通过张量补全优化的CP分解在高维应用场景中具有更高的预测精度与内存效率。