Foundation models are becoming the dominant deep learning technologies. Pretraining a foundation model is always time-consumed due to the large scale of both the model parameter and training dataset. Besides being computing-intensive, the training process is extremely memory-intensive and communication-intensive. These features make it necessary to apply 3D parallelism, which integrates data parallelism, pipeline model parallelism and tensor model parallelism, to achieve high training efficiency. To achieve this goal, some custom software frameworks such as Megatron-LM and DeepSpeed are developed. However, current 3D parallelism frameworks still meet two issues: i) they are not transparent to model developers, which need to manually modify the model to parallelize training. ii) their utilization of computation, GPU memory and network bandwidth are not sufficient. We propose Merak, an automated 3D parallelism deep learning training framework with high resource utilization. Merak automatically deploys with an automatic model partitioner, which uses a graph sharding algorithm on a proxy representation of the model. Merak also presents the non-intrusive API for scaling out foundation model training with minimal code modification. In addition, we design a high-performance 3D parallel runtime engine in Merak. It uses several techniques to exploit available training resources, including shifted critical path pipeline schedule that brings a higher computation utilization, stage-aware recomputation that makes use of idle worker memory, and sub-pipelined tensor model parallelism that overlaps communication and computation. Experiments on 64 GPUs show Merak can speedup the training performance over the state-of-the-art 3D parallelism frameworks of models with 1.5, 2.5, 8.3, and 20 billion parameters by up to 1.42X, 1.39X, 1.43X, and 1.61X, respectively.


翻译:基础模型正成为主流的深度学习技术。由于模型参数和训练数据集的规模庞大,预训练基础模型往往耗时巨大。除计算密集外,训练过程还极度消耗内存和通信资源。这些特性使得必须采用三维并行(集成数据并行、流水线模型并行和张量模型并行)以实现高训练效率。为此,业界开发了诸如Megatron-LM和DeepSpeed等定制化软件框架。然而,当前的三维并行框架仍存在两个问题:i) 对模型开发者不透明,需手动修改模型以实现并行训练;ii) 对计算资源、GPU内存和网络带宽的利用率不足。我们提出Merak——一种高资源利用率的自动化三维并行深度学习训练框架。Merak通过自动模型分割器实现自动部署,该分割器在模型的代理表示上采用图分片算法。Merak还提供了非侵入式API,仅需极少量代码修改即可扩展基础模型训练。此外,我们设计了Merak的高性能三维并行运行时引擎,采用多项技术以充分利用可用训练资源,包括:通过偏移关键路径流水线调度提升计算利用率;利用空闲工作节点内存的阶段感知重计算;以及支持通信与计算重叠的子流水线张量模型并行。在64块GPU上的实验表明,针对参数规模为1.5亿、2.5亿、8.3亿和200亿的模型,Merak相较于当前最先进的三维并行框架,训练性能分别提升至1.42倍、1.39倍、1.43倍和1.61倍。

0
下载
关闭预览

相关内容

ACM/IEEE第23届模型驱动工程语言和系统国际会议,是模型驱动软件和系统工程的首要会议系列,由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来,模型涵盖了建模的各个方面,从语言和方法到工具和应用程序。模特的参加者来自不同的背景,包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛,参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会,并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。 官网链接:http://www.modelsconference.org/
【普林斯顿博士论文】构建高效深度神经网络,195页pdf
专知会员服务
70+阅读 · 2023年2月8日
【伯克利博士论文】硬件感知的高效深度学习,154页pdf
专知会员服务
75+阅读 · 2022年10月20日
【伯克利Alvin Wan博士论文】高效设计深度神经网络
专知会员服务
61+阅读 · 2022年5月21日
【2022新书】高效深度学习,Efficient Deep Learning Book
专知会员服务
128+阅读 · 2022年4月21日
实践教程|PyTorch 并行训练极简 Demo
极市平台
0+阅读 · 2022年11月12日
TF Boys必看!一文搞懂TensorFlow 2.0新架构!
引力空间站
19+阅读 · 2019年1月16日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
1+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
1+阅读 · 2012年12月31日
国家自然科学基金
8+阅读 · 2009年12月31日
国家自然科学基金
2+阅读 · 2009年12月31日
Arxiv
0+阅读 · 2023年5月11日
Arxiv
20+阅读 · 2019年11月23日
VIP会员
最新内容
《基于强化学习的自动化红队测试》
专知会员服务
3+阅读 · 7月23日
伊朗不对称防空战略的演进
专知会员服务
4+阅读 · 7月23日
对抗环境下超视距目标打击的情报支援
专知会员服务
10+阅读 · 7月22日
《无人机对海面作战影响评估》
专知会员服务
15+阅读 · 7月21日
印度精确打击与指挥架构的断层
专知会员服务
7+阅读 · 7月20日
相关VIP内容
【普林斯顿博士论文】构建高效深度神经网络,195页pdf
专知会员服务
70+阅读 · 2023年2月8日
【伯克利博士论文】硬件感知的高效深度学习,154页pdf
专知会员服务
75+阅读 · 2022年10月20日
【伯克利Alvin Wan博士论文】高效设计深度神经网络
专知会员服务
61+阅读 · 2022年5月21日
【2022新书】高效深度学习,Efficient Deep Learning Book
专知会员服务
128+阅读 · 2022年4月21日
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
1+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
1+阅读 · 2012年12月31日
国家自然科学基金
8+阅读 · 2009年12月31日
国家自然科学基金
2+阅读 · 2009年12月31日
Top
微信扫码咨询专知VIP会员