The ever-increasing computational and storage requirements of modern applications and the slowdown of technology scaling pose major challenges to designing and implementing efficient computer architectures. In this paper, we leverage the architectural balance principle to alleviate the bandwidth bottleneck at the L1 data memory boundary of a tightly-coupled cluster of processing elements (PEs). We thus explore coupling each PE with an L0 memory, namely a private register file implemented as Standard Cell Memory (SCM). Architecturally, the SCM is the Vector Register File (VRF) of Spatz, a compact 64-bit floating-point-capable vector processor based on RISC-V's Vector Extension Zve64d. Unlike typical vector processors, whose VRF are hundreds of KiB large, we prove that Spatz can achieve peak energy efficiency with a VRF of only 2 KiB. An implementation of the Spatz-based cluster in GlobalFoundries' 12LPP process with eight double-precision Floating Point Units (FPUs) achieves an FPU utilization just 3.4% lower than the ideal upper bound on a double-precision, floating-point matrix multiplication. The cluster reaches 7.7 FMA/cycle, corresponding to 15.7 GFLOPS-DP and 95.7 GFLOPS-DP/W at 1 GHz and nominal operating conditions (TT, 0.80V, 25^oC) with more than 55% of the power spent on the FPUs. Furthermore, the optimally-balanced Spatz-based cluster reaches a 95.0% FPU utilization (7.6 FMA/cycle), 15.2 GFLOPS-DP, and 99.3 GFLOPS-DP/W (61% of the power spent in the FPU) on a 2D workload with a 7x7 kernel, resulting in an outstanding area/energy efficiency of 171 GFLOPS-DP/W/mm^2. At equi-area, our computing cluster built upon compact vector processors reaches a 30% higher energy efficiency than a cluster with the same FPU count built upon scalar cores specialized for stream-based floating-point computation.
翻译:现代应用对计算和存储需求的不断增长,以及技术缩放速度的放缓,给设计和实现高效计算机架构带来了重大挑战。本文利用架构平衡原则,缓解处理单元紧密耦合集群中L1数据存储器边界的带宽瓶颈。为此,我们探索为每个PE配备L0存储器,即实现为标准单元存储器(SCM)的私有寄存器文件。架构上,该SCM即Spatz(一款基于RISC-V向量扩展Zve64d的紧凑型64位浮点向量处理器)的向量寄存器文件(VRF)。与典型向量处理器(其VRF通常达数百KiB)不同,我们证明Spatz仅需2 KiB的VRF即可达到峰值能效。基于Spatz的集群在GlobalFoundries 12LPP工艺中实现,配备八个双精度浮点运算单元,在双精度浮点矩阵乘法中,FPU利用率仅比理想上限低3.4%。该集群在1 GHz标称工作条件(TT, 0.80V, 25°C)下达到7.7 FMA/周期,对应15.7 GFLOPS-DP和95.7 GFLOPS-DP/W,其中超过55%的功耗用于FPU。此外,最优平衡的Spatz集群在7x7核的2D工作负载上达到95.0%的FPU利用率(7.6 FMA/周期)、15.2 GFLOPS-DP和99.3 GFLOPS-DP/W(61%功耗用于FPU),实现了卓越的面积/能效比171 GFLOPS-DP/W/mm²。在同等面积下,我们的紧凑型向量处理器集群相比采用相同FPU数量、专用于流式浮点计算的标量核心集群,能效提升了30%。