Alternating Updates for Efficient Transformers - 专知论文

会员服务 ·

0

变换 · MoDELS · 缩放 · 计算成本 · 代价 ·

2023 年 1 月 30 日

Alternating Updates for Efficient Transformers

翻译：面向高效Transformer的交替更新方法

Cenk Baykal,Dylan Cutler,Nishanth Dikkala,Nikhil Ghosh,Rina Panigrahy,Xin Wang

It is well established that increasing scale in deep transformer networks leads to improved quality and performance. This increase in scale often comes with an increase in compute cost and inference latency. Consequently, research into methods which help realize the benefits of increased scale without leading to an increase in the compute cost becomes important. We introduce Alternating Updates (AltUp), a simple-to-implement method to increase a model's capacity without the computational burden. AltUp enables the widening of the learned representation without increasing the computation time by working on a subblock of the representation at each layer. Our experiments on various transformer models and language tasks demonstrate the consistent effectiveness of alternating updates on a diverse set of benchmarks. Finally, we present extensions of AltUp to the sequence dimension, and demonstrate how AltUp can be synergistically combined with existing approaches, such as Sparse Mixture-of-Experts models, to obtain efficient models with even higher capacity.

翻译：众所周知，深度Transformer网络规模的扩大会带来模型质量与性能的提升。然而规模扩大往往伴随计算成本增加与推理延迟升高。因此，研究如何在避免计算成本增长的前提下实现规模扩大带来的效益变得至关重要。本文提出交替更新（AltUp）——一种易于实现的方法，能在不增加计算负担的情况下提升模型容量。AltUp通过每层处理表示的子块，在不增加计算时间的前提下扩展学习表示的宽度。我们在多种Transformer模型和语言任务上的实验表明，交替更新在各类基准测试中均保持稳定有效性。最后，我们将AltUp扩展至序列维度，并展示其如何与现有方法（如稀疏混合专家模型）协同结合，从而获得容量更高且计算高效的模型。

0

相关内容

NeurlPS 2022 | 自然语言处理相关论文分类整理

NeurlPS 2022 | 自然语言处理相关论文分类整理

专知会员服务

51+阅读 · 2022年10月2日

【2022新书】高效深度学习，Efficient Deep Learning Book

【2022新书】高效深度学习，Efficient Deep Learning Book

专知会员服务

128+阅读 · 2022年4月21日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

剑桥大学《数据科学: 原理与实践》课程，附PPT下载

剑桥大学《数据科学: 原理与实践》课程，附PPT下载

专知会员服务

54+阅读 · 2021年1月20日

最新《Transformers模型》教程，64页ppt

最新《Transformers模型》教程，64页ppt

专知会员服务

326+阅读 · 2020年11月26日

UC.Berkeley CS189讲义教材:《机器学习全面指南》，185页pdf

专知会员服务

162+阅读 · 2020年1月16日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

44+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

MAI/H2O2/NIR三重应答脂质体改善“缺氧”肿瘤的PDT效果的研究

国家自然科学基金

0+阅读 · 2015年12月31日

微流控芯片可控合成siRNA复合纳米材料及其应用

国家自然科学基金

0+阅读 · 2014年12月31日

激光二极管触发GaAs光导开关的失效机理研究

国家自然科学基金

0+阅读 · 2013年12月31日

大规模混合设计变量结构优化设计研究

国家自然科学基金

0+阅读 · 2013年12月31日

DNA甲基化和去甲基化在神经干细胞向星形胶质细胞分化过程中的作用机制

国家自然科学基金

0+阅读 · 2012年12月31日

求解对流扩散方程的全离散间断有限元方法

国家自然科学基金

0+阅读 · 2012年12月31日

高活性磷化镍催化剂的设计制备及其深度加氢脱硫活性调控机制

国家自然科学基金

0+阅读 · 2012年12月31日

Dicer在慢性乙型病毒性肝炎恶性转化过程中的作用

国家自然科学基金

0+阅读 · 2011年12月31日

能量临界情形的非线性Schrodinger方程

国家自然科学基金

0+阅读 · 2011年12月31日

病毒感染中p100蛋白与I-dsRNA相互作用的分子机制研究

国家自然科学基金

0+阅读 · 2011年12月31日

μSplit: efficient image decomposition for microscopy data

Arxiv

0+阅读 · 2023年3月22日

On Penalty-based Bilevel Gradient Descent Method

Arxiv

0+阅读 · 2023年3月21日

eP-ALM: Efficient Perceptual Augmentation of Language Models

Arxiv

0+阅读 · 2023年3月20日

Robustifying Token Attention for Vision Transformers

Arxiv

0+阅读 · 2023年3月20日

N-Gram in Swin Transformers for Efficient Lightweight Image Super-Resolution

Arxiv

0+阅读 · 2023年3月20日

Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning

Arxiv

1+阅读 · 2023年3月18日

Treeformer: Dense Gradient Trees for Efficient Attention Computation

Arxiv

0+阅读 · 2023年3月17日

SFE: A Simple, Fast and Efficient Feature Selection Algorithm for High-Dimensional Data

Arxiv

0+阅读 · 2023年3月17日

Over-Fit: Noisy-Label Detection based on the Overfitted Model Property

Arxiv

0+阅读 · 2023年3月17日

Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting

Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting

Arxiv

21+阅读 · 2020年12月17日

VIP会员

文章信息

相关主题

最新内容

《无人机对海面作战影响评估》

《无人机对海面作战影响评估》

专知会员服务

9+阅读 · 7月21日

《可损耗无人系统规模化应用对美国军事转型的战略影响（2022-2030）》2026年270页

《可损耗无人系统规模化应用对美国军事转型的战略影响（2022-2030）》2026年270页

专知会员服务

8+阅读 · 7月21日

博士论文 | 后训练如何损害大模型生成多样性？SimpleStrat与Stylus

博士论文 | 后训练如何损害大模型生成多样性？SimpleStrat与Stylus

专知会员服务

3+阅读 · 7月21日

综述 | 面向5G/6G网络的LLM智能体AI：架构、协议与标准化

综述 | 面向5G/6G网络的LLM智能体AI：架构、协议与标准化

专知会员服务

5+阅读 · 7月21日

五角大楼新设无人机办公室（DRPM-UxS）将如何重塑美国无人系统格局（附美国防部设立备忘录）

五角大楼新设无人机办公室（DRPM-UxS）将如何重塑美国无人系统格局（附美国防部设立备忘录）

专知会员服务

6+阅读 · 7月21日

印度精确打击与指挥架构的断层

印度精确打击与指挥架构的断层

专知会员服务

5+阅读 · 7月20日

《NASA喷气推进实验室：高耐久轻质常驻空观测系统（HELIOS）》429页

《NASA喷气推进实验室：高耐久轻质常驻空观测系统（HELIOS）》429页

专知会员服务

7+阅读 · 7月20日

美空军AI完成F-16战斗机自主空战历史性试飞

美空军AI完成F-16战斗机自主空战历史性试飞

专知会员服务

6+阅读 · 7月20日

《美政府问责局——武器系统年度评估（2026年）：强制要求成熟技术或可推动转向快速交付》249页

《美政府问责局——武器系统年度评估（2026年）：强制要求成熟技术或可推动转向快速交付》249页

专知会员服务

8+阅读 · 7月20日

《美国陆军：通过弹性分布式模型库实现自适应AI优势》

《美国陆军：通过弹性分布式模型库实现自适应AI优势》

专知会员服务

7+阅读 · 7月20日

博士论文 | 理解与改进大语言模型推理：从反转诅咒到连续思维链

博士论文 | 理解与改进大语言模型推理：从反转诅咒到连续思维链

专知会员服务

9+阅读 · 7月20日

综述 | 终身视觉表征：持续自监督学习CSSL系统综述

综述 | 终身视觉表征：持续自监督学习CSSL系统综述

专知会员服务

9+阅读 · 7月20日

深入Project Maven：为何人工智能在战场上依然失灵

深入Project Maven：为何人工智能在战场上依然失灵

专知会员服务

15+阅读 · 7月19日

锻造未来士兵：外骨骼、基因工程与赛博格

锻造未来士兵：外骨骼、基因工程与赛博格

专知会员服务

8+阅读 · 7月19日

《无人机系统（UAS）通信网状网络试验性部署》50页报告

《无人机系统（UAS）通信网状网络试验性部署》50页报告

专知会员服务

10+阅读 · 7月19日

相关VIP内容

NeurlPS 2022 | 自然语言处理相关论文分类整理

NeurlPS 2022 | 自然语言处理相关论文分类整理

专知会员服务

51+阅读 · 2022年10月2日

【2022新书】高效深度学习，Efficient Deep Learning Book

【2022新书】高效深度学习，Efficient Deep Learning Book

专知会员服务

128+阅读 · 2022年4月21日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

剑桥大学《数据科学: 原理与实践》课程，附PPT下载

剑桥大学《数据科学: 原理与实践》课程，附PPT下载

专知会员服务

54+阅读 · 2021年1月20日

最新《Transformers模型》教程，64页ppt

最新《Transformers模型》教程，64页ppt

专知会员服务

326+阅读 · 2020年11月26日

UC.Berkeley CS189讲义教材:《机器学习全面指南》，185页pdf

专知会员服务

162+阅读 · 2020年1月16日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

《可损耗无人系统规模化应用对美国军事转型的战略影响（2022-2030）》2026年270页

综述 | 面向5G/6G网络的LLM智能体AI：架构、协议与标准化

《无人机对海面作战影响评估》

博士论文 | 后训练如何损害大模型生成多样性？SimpleStrat与Stylus

相关资讯

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

44+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

相关论文

μSplit: efficient image decomposition for microscopy data

Arxiv

0+阅读 · 2023年3月22日

On Penalty-based Bilevel Gradient Descent Method

Arxiv

0+阅读 · 2023年3月21日

eP-ALM: Efficient Perceptual Augmentation of Language Models

Arxiv

0+阅读 · 2023年3月20日

Robustifying Token Attention for Vision Transformers

Arxiv

0+阅读 · 2023年3月20日

N-Gram in Swin Transformers for Efficient Lightweight Image Super-Resolution

Arxiv

0+阅读 · 2023年3月20日

Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning

Arxiv

1+阅读 · 2023年3月18日

Treeformer: Dense Gradient Trees for Efficient Attention Computation

Arxiv

0+阅读 · 2023年3月17日

SFE: A Simple, Fast and Efficient Feature Selection Algorithm for High-Dimensional Data

Arxiv

0+阅读 · 2023年3月17日

Over-Fit: Noisy-Label Detection based on the Overfitted Model Property

Arxiv

0+阅读 · 2023年3月17日

Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting

Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting

Arxiv

21+阅读 · 2020年12月17日

相关基金

MAI/H2O2/NIR三重应答脂质体改善“缺氧”肿瘤的PDT效果的研究

国家自然科学基金

0+阅读 · 2015年12月31日

微流控芯片可控合成siRNA复合纳米材料及其应用

国家自然科学基金

0+阅读 · 2014年12月31日

激光二极管触发GaAs光导开关的失效机理研究

国家自然科学基金

0+阅读 · 2013年12月31日

大规模混合设计变量结构优化设计研究

国家自然科学基金

0+阅读 · 2013年12月31日

DNA甲基化和去甲基化在神经干细胞向星形胶质细胞分化过程中的作用机制

国家自然科学基金

0+阅读 · 2012年12月31日

求解对流扩散方程的全离散间断有限元方法

国家自然科学基金

0+阅读 · 2012年12月31日

高活性磷化镍催化剂的设计制备及其深度加氢脱硫活性调控机制

国家自然科学基金

0+阅读 · 2012年12月31日

Dicer在慢性乙型病毒性肝炎恶性转化过程中的作用

国家自然科学基金

0+阅读 · 2011年12月31日

能量临界情形的非线性Schrodinger方程

国家自然科学基金

0+阅读 · 2011年12月31日

病毒感染中p100蛋白与I-dsRNA相互作用的分子机制研究

国家自然科学基金

0+阅读 · 2011年12月31日

微信扫码咨询专知VIP会员