We present thread-safe, highly-optimized lattice Boltzmann implementations, specifically aimed at exploiting the high memory bandwidth of GPU-based architectures. At variance with standard approaches to LB coding, the proposed strategy, based on the reconstruction of the post-collision distribution via Hermite projection, enforces data locality and avoids the onset of memory dependencies, which may arise during the propagation step, with no need to resort to more complex streaming strategies. The thread-safe lattice Boltzmann achieves peak performances, both in two and three dimensions and it allows to sensibly reduce the allocated memory ( tens of GigaBytes for order billions lattice nodes simulations) by retaining the algorithmic simplicity of standard LB computing. Our findings open attractive prospects for high-performance simulations of complex flows on GPU-based architectures.
翻译:我们提出了线程安全且高度优化的格子玻尔兹曼实现方案,旨在充分利用基于GPU架构的高内存带宽。与标准LB编码方法不同,本策略基于赫米特投影重构碰撞后分布,强制执行数据局部性并避免传播步骤中可能出现的内存依赖问题,无需采用更复杂的流策略。该线程安全格子玻尔兹曼方法在二维和三维场景中均能达到峰值性能,并通过保留标准LB计算的算法简洁性,显著减少内存分配(对数十亿节点量级的模拟可节约数十GB内存)。我们的研究成果为基于GPU架构的复杂流高性能模拟开辟了极具吸引力的前景。