Graphics processing units (GPUs) are continually evolving to cater to the computational demands of contemporary general-purpose workloads, particularly those driven by artificial intelligence (AI) utilizing deep learning techniques. A substantial body of studies have been dedicated to dissecting the microarchitectural metrics characterizing diverse GPU generations, which helps researchers understand the hardware details and leverage them to optimize the GPU programs. However, the latest Hopper GPUs present a set of novel attributes, including new tensor cores supporting FP8, DPX, and distributed shared memory. Their details still remain mysterious in terms of performance and operational characteristics. In this research, we propose an extensive benchmarking study focused on the Hopper GPU. The objective is to unveil its microarchitectural intricacies through an examination of the new instruction-set architecture (ISA) of Nvidia GPUs and the utilization of new CUDA APIs. Our approach involves two main aspects. Firstly, we conduct conventional latency and throughput comparison benchmarks across the three most recent GPU architectures, namely Hopper, Ada, and Ampere. Secondly, we delve into a comprehensive discussion and benchmarking of the latest Hopper features, encompassing the Hopper DPX dynamic programming (DP) instruction set, distributed shared memory, and the availability of FP8 tensor cores. The microbenchmarking results we present offer a deeper understanding of the novel GPU AI function units and programming features introduced by the Hopper architecture. This newfound understanding is expected to greatly facilitate software optimization and modeling efforts for GPU architectures. To the best of our knowledge, this study makes the first attempt to demystify the tensor core performance and programming instruction sets unique to Hopper GPUs.
翻译:图形处理器(GPU)正持续演进以满足当代通用计算工作负载的需求,尤其是基于深度学习技术的人工智能(AI)应用。大量研究致力于剖析不同代际GPU的微架构指标,这有助于研究者理解硬件细节并利用其优化GPU程序。然而,最新的Hopper GPU展现出多项新特性,包括支持FP8、DPX的新型张量核心以及分布式共享内存,其性能和运行特性的细节仍充满神秘色彩。本研究针对Hopper GPU开展了详尽的基准测试,旨在通过分析Nvidia GPU的新指令集架构(ISA)及新CUDA API的使用来揭示其微架构细节。我们的方法包含两个方面:首先,我们对Hopper、Ada和Ampere这三款最新GPU架构进行了常规延迟与吞吐量对比基准测试;其次,我们深入探讨并测试了Hopper的最新特性,包括Hopper DPX动态规划(DP)指令集、分布式共享内存以及FP8张量核心的可用性。本文呈现的微基准测试结果有助于深入理解Hopper架构引入的新型GPU AI功能单元与编程特性。这一新认知将极大促进GPU架构的软件优化与建模工作。据我们所知,本研究首次尝试揭秘Hopper GPU独有的张量核心性能与编程指令集。