In recent years, Transformer-based Large Language Models (LLMs) have garnered significant attention due to their exceptional performance across a variety of tasks. However, training these models on long sequences presents a substantial challenge in terms of efficiency and scalability. Current methods are constrained either by the number of attention heads, limiting scalability, or by excessive communication overheads. In this paper, we propose an insight that Attention Computation can be considered as a special case of n-body problem with direct interactions. Based on this concept, this paper introduces WallFacer, an efficient long-sequence training system with a novel multi-dimensional ring sequence parallelism, fostering an efficient communication paradigm and extra tuning space for communication arrangement. Through comprehensive experiments under diverse environments and model settings, we demonstrate that WallFacer significantly surpasses state-of-the-art method that supports near-infinite sequence length, achieving performance improvements of up to 77.12%.
翻译:近年来,基于Transformer的大语言模型(LLMs)因其在多种任务上的卓越性能而受到广泛关注。然而,在长序列上训练这些模型在效率和可扩展性方面存在重大挑战。现有方法要么受限于注意力头数量而制约可扩展性,要么面临过高的通信开销。本文提出一种洞见:注意力计算可被视为具有直接相互作用的N体问题的一种特例。基于这一概念,本文提出了WallFacer——一种高效的长序列训练系统,其采用新颖的多维环形序列并行策略,构建了高效的通信范式并为通信调度提供了额外的优化空间。通过在多样化环境和模型设置下的综合实验,我们证明WallFacer显著超越了支持近无限序列长度的最先进方法,性能提升最高可达77.12%。