The recent trend toward deep learning has led to the development of a variety of highly innovative AI accelerator architectures. One such architecture, the Cerebras Wafer-Scale Engine 2 (WSE-2), features 40 GB of on-chip SRAM, making it a potentially attractive platform for latency- or bandwidth-bound HPC simulation workloads. In this study, we examine the feasibility of performing continuous energy Monte Carlo (MC) particle transport on the WSE-2 by porting a key kernel from the MC transport algorithm to Cerebras's CSL programming model. New algorithms for minimizing communication costs and for handling load balancing are developed and tested. The WSE-2 is found to run \SPEEDUP~times faster than a highly optimized CUDA version of the kernel run on an NVIDIA A100 GPU -- significantly outpacing the expected performance increase given the difference in transistor counts between the architectures.
翻译:深度学习的最新趋势促进了多种极具创新性的AI加速器架构的发展。其中,Cerebras晶圆级引擎2(WSE-2)配备了40 GB的片上SRAM,使其成为延迟敏感或带宽受限的高性能计算模拟任务中极具吸引力的平台。本研究通过将蒙特卡洛(MC)输运算法中的关键内核移植至Cerebras的CSL编程模型,探讨了在WSE-2上执行连续能量蒙特卡洛粒子输运的可行性。我们开发并测试了用于最小化通信开销和实现负载均衡的新算法。实验发现,与高度优化的NVIDIA A100 GPU上的CUDA版本内核相比,WSE-2的运行速度提升\SPEEDUP~倍——这显著超越了基于两种架构晶体管数量差异所预期的性能增益。