Transformers have revolutionized deep learning and generative modeling, enabling unprecedented advancements in natural language processing tasks. However, the size of transformer models is increasing continuously, driven by enhanced capabilities across various deep-learning tasks. This trend of ever-increasing model size has given rise to new challenges in terms of memory and computing requirements. Conventional computing platforms, including GPUs, suffer from suboptimal performance due to the memory demands imposed by models with millions/billions of parameters. The emerging chiplet-based platforms provide a new avenue for compute- and data-intensive machine learning (ML) applications enabled by a Network-on-Interposer (NoI). However, designing suitable hardware accelerators for executing Transformer inference workloads is challenging due to a wide variety of complex computing kernels in the Transformer architecture. In this paper, we leverage chiplet-based heterogeneous integration (HI) to design a high-performance and energy-efficient multi-chiplet platform to accelerate transformer workloads. We demonstrate that the proposed NoI architecture caters to the data access patterns inherent in a transformer model. The optimized placement of the chiplets and the associated NoI links and routers enable superior performance compared to the state-of-the-art hardware accelerators. The proposed NoI-based architecture demonstrates scalability across varying transformer models and improves latency and energy efficiency by up to 22.8x and 5.36x respectively.
翻译:Transformer模型已彻底改变了深度学习与生成式建模领域,推动了自然语言处理任务取得前所未有的进步。然而,随着模型在各类深度学习任务中能力的持续增强,其规模也不断扩大。这一模型规模持续增长的趋势带来了内存与计算需求方面的全新挑战。传统计算平台(包括GPU)在处理具有数百万/数十亿参数的模型时,因内存需求受限而表现欠佳。新兴的基于中介层网络(Network-on-Interposer, NoI)的芯粒平台,为计算密集型和数据密集型的机器学习应用提供了新路径。然而,由于Transformer架构中包含多种复杂的计算核心,设计适用于执行Transformer推理工作负载的硬件加速器仍面临诸多挑战。本文利用基于芯粒的异构集成(Heterogeneous Integration, HI)技术,设计了一种高性能、高能效的多芯粒平台,用于加速Transformer工作负载。我们证明了所提出的NoI架构能够适配Transformer模型固有的数据访问模式。通过优化芯粒布局、NoI链路与路由器设计,该架构相较于当前最先进的硬件加速器展现出更优异的性能。基于NoI的架构在不同Transformer模型间具有良好的可扩展性,且延迟和能效分别最高提升了22.8倍和5.36倍。