Ootomo, Ozaki, and Yokota [Int. J. High Perform. Comput. Appl., 38 (2024), p. 297-313] have proposed a strategy to recast a floating-point matrix multiplication in terms of integer matrix products. The factors A and B are split into integer slices, the product of these slices is computed exactly, and AB is approximated by accumulating these integer products in floating-point arithmetic. This technique is particularly well suited to mixed-precision matrix multiply-accumulate units with integer support, such as the NVIDIA tensor cores or the AMD matrix cores. The number of slices allows for performance-accuracy tradeoffs: more slices yield better accuracy but require more multiplications, which in turn reduce performance. We propose an inexpensive way to estimate the minimum number of multiplications needed to achieve a prescribed level of accuracy. Our error analysis shows that the algorithm may become inaccurate (or inefficient) if rows of A or columns of B are badly scaled. We perform a range of numerical experiments, both in simulation and on the latest NVIDIA GPUs, that confirm the analysis and illustrate strengths and weaknesses of the algorithm.
翻译:Ootomo、Ozaki和Yokota [Int. J. High Perform. Comput. Appl., 38 (2024), p. 297-313] 提出了一种策略,将浮点矩阵乘法转化为整数矩阵乘积。因子A和B被分割成整数切片,这些切片的乘积被精确计算,通过将这些整数乘积在浮点算术中累加,得到AB的近似值。该技术特别适用于支持整数的混合精度矩阵乘累加单元,如NVIDIA张量核心或AMD矩阵核心。切片数量允许在性能与精度之间进行权衡:更多切片可提高精度,但需要更多乘法运算,从而降低性能。我们提出了一种低开销的方法,用于估计达到指定精度水平所需的最小乘法次数。我们的误差分析表明,若A的行或B的列存在严重的尺度不平衡,该算法可能变得不精确(或低效)。我们进行了一系列数值实验,包括仿真和在最新NVIDIA GPU上的实验,验证了分析结果,并展示了该算法的优缺点。