Challenging the Nvidia monopoly, dedicated AI-accelerator chips have begun emerging for tackling the computational challenge that the inference and, especially, the training of modern deep neural networks (DNNs) poses to modern computers. The field has been ridden with studies assessing the performance of these contestants across various DNN model types. However, AI-experts are aware of the limitations of current DNNs and have been working towards the fourth AI wave which will, arguably, rely on more biologically inspired models, predominantly on spiking neural networks (SNNs). At the same time, GPUs have been heavily used for simulating such models in the field of computational neuroscience, yet AI-chips have not been tested on such workloads. The current paper aims at filling this important gap by evaluating multiple, cutting-edge AI-chips (Graphcore IPU, GroqChip, Nvidia GPU with Tensor Cores and Google TPU) on simulating a highly biologically detailed model of a brain region, the inferior olive (IO). This IO application stress-tests the different AI-platforms for highlighting architectural tradeoffs by varying its compute density, memory requirements and floating-point numerical accuracy. Our performance analysis reveals that the simulation problem maps extremely well onto the GPU and TPU architectures, which for networks of 125,000 cells leads to a 28x respectively 1,208x speedup over CPU runtimes. At this speed, the TPU sets a new record for largest real-time IO simulation. The GroqChip outperforms both platforms for small networks but, due to implementing some floating-point operations at reduced accuracy, is found not yet usable for brain simulation.
翻译:挑战英伟达的垄断地位,专用AI加速器芯片开始涌现,以应对现代深度神经网络(DNN)推理,尤其是训练给现代计算机带来的计算挑战。该领域已有大量研究评估这些竞争者在各类DNN模型类型上的性能表现。然而,AI专家已意识到当前DNN的局限性,并正致力于推动第四次AI浪潮——这一浪潮可能将主要依赖更具生物启发性的模型,尤其是脉冲神经网络(SNN)。与此同时,GPU已被广泛用于计算神经科学领域模拟此类模型,但AI芯片尚未在类似工作负载下接受测试。本文旨在填补这一重要空白,通过评估多款前沿AI芯片(Graphcore IPU、GroqChip、配备张量核心的英伟达GPU及谷歌TPU)在模拟高度生物细节的大脑区域——下橄榄核(IO)模型时的表现。该IO应用通过改变计算密度、内存需求及浮点数数值精度,对不同AI平台进行压力测试,以突显架构权衡。我们的性能分析表明,该模拟问题与GPU和TPU架构高度匹配:在12.5万个细胞的网络中,两者相比CPU运行时间分别实现了28倍和1208倍的加速。在此速度下,TPU创下了最大规模实时IO模拟的新纪录。GroqChip在小型网络中性能优于上述两个平台,但由于部分浮点运算采用降精度实现,目前尚不适用于大脑模拟。