The expanding size of language models has created the necessity for a comprehensive examination across various dimensions that reflect the desiderata with respect to the tradeoffs between various hardware metrics, such as latency, energy consumption, GPU memory usage, and performance. There is a growing interest in establishing Pareto frontiers for different language model configurations to identify optimal models with specified hardware constraints. Notably, architectures that excel in latency on one device may not perform optimally on another. However, exhaustive training and evaluation of numerous architectures across diverse hardware configurations is computationally prohibitive. To this end, we propose HW-GPT-Bench, a hardware-aware language model surrogate benchmark, where we leverage weight-sharing techniques from Neural Architecture Search (NAS) to efficiently train a supernet proxy, encompassing language models of varying scales in a single model. We conduct profiling of these models across 13 devices, considering 5 hardware metrics and 3 distinct model scales. Finally, we showcase the usability of HW-GPT-Bench using 8 different multi-objective NAS algorithms and evaluate the quality of the resultant Pareto fronts. Through this benchmark, our objective is to propel and expedite research in the advancement of multi-objective methods for NAS and structural pruning in large language models.
翻译:随着语言模型规模的不断扩展,亟需从多个维度对其展开全面研究,以反映延迟、能耗、GPU内存使用与性能等不同硬件指标之间的权衡需求。针对不同语言模型配置建立帕累托前沿,以识别在特定硬件约束下的最优模型,正引起学界日益浓厚的兴趣。值得注意的是,在某类设备上延迟表现优异的架构,在另一类设备上未必能达到最优。然而,在多样化的硬件配置下对大量架构进行穷举式训练与评估,其计算成本高得令人望而却步。为此,我们提出HW-GPT-Bench——一种硬件感知的语言模型替代基准。该方法利用神经架构搜索(NAS)中的权重共享技术,高效训练一个超网络代理,将不同规模的语言模型整合于单一模型之中。我们在13种设备上对这些模型进行性能剖析,覆盖5项硬件指标及3种不同模型规模。最后,我们通过8种不同的多目标NAS算法展示了HW-GPT-Bench的实用性,并评估了所得帕累托前沿的质量。通过该基准,我们旨在推动并加速面向大型语言模型中NAS及结构剪枝的多目标方法研究进展。