Parcae: Scaling Laws For Stable Looped Language Models - 专知论文

会员服务 ·

0

Parcae: Scaling Laws For Stable Looped Language Models

翻译：暂无翻译

Hayden Prairie,Zachary Novack,Taylor Berg-Kirkpatrick,Daniel Y. Fu

Traditional fixed-depth architectures scale quality by increasing training FLOPs, typically through increased parameterization, at the expense of a higher memory footprint, or data. A potential alternative is looped architectures, which instead increase FLOPs by sending activations through a block of layers in a loop. While promising, existing recipes for training looped architectures can be unstable, suffering from residual explosion and loss spikes. We address these challenges by recasting looping as a nonlinear time-variant dynamical system over the residual stream. Via a linear approximation to this system, we find that instability occurs in existing looped architectures as a result of large spectral norms in their injection parameters. To address these instability issues, we propose Parcae, a novel stable, looped architecture that constrains the spectral norm of the injection parameters via discretization of a negative diagonal parameterization. As a result, Parcae achieves up to 6.3% lower validation perplexity over prior large-scale looped models. Using our stable looped architecture, we investigate the scaling properties of looping as a medium to improve quality by increasing FLOPs in training and test-time. For training, we derive predictable power laws to scale FLOPs while keeping parameter count fixed. Our initial scaling laws suggest that looping and data should be increased in tandem, given a fixed FLOP budget. At test-time, we find that Parcae can use looping to scale compute, following a predictable, saturating exponential decay. When scaled up to 1.3B parameters, we find that Parcae improves CORE and Core-Extended quality by 2.99 and 1.18 points when compared to strong Transformer baselines under a fixed parameter and data budget, achieving a relative quality of up to 87.5% a Transformer twice the size.

翻译：暂无翻译

0

相关内容

不可错过！EPFL《训练大语言模型》课程

不可错过！EPFL《训练大语言模型》课程

专知会员服务

18+阅读 · 2025年4月25日

最高9.0分！这16篇最高分ICLR2025论文必看！从生成模型到MOE等

最高9.0分！这16篇最高分ICLR2025论文必看！从生成模型到MOE等

专知会员服务

26+阅读 · 2024年11月19日

视频大模型中视觉上下文表示的scaling law

视频大模型中视觉上下文表示的scaling law

专知会员服务

24+阅读 · 2024年10月21日

【TPAMI2024】结构化时空对齐视频-语言表示

【TPAMI2024】结构化时空对齐视频-语言表示

专知会员服务

29+阅读 · 2024年10月20日

ACL2024 | IEPILE:大规模基于Schema的信息抽取语料库

ACL2024 | IEPILE:大规模基于Schema的信息抽取语料库

专知会员服务

32+阅读 · 2024年6月20日

国家标准《人工智能预训练模型第3 部分服务能力成熟度评估》

国家标准《人工智能预训练模型第3 部分服务能力成熟度评估》

专知会员服务

63+阅读 · 2024年6月16日

【CVPR 2022】连续驾驶场景与不断增长的建筑的连续立体匹配，Continual Stereo Matching of Continuous Driving Scenes with Growing Architecture

【CVPR 2022】连续驾驶场景与不断增长的建筑的连续立体匹配，Continual Stereo Matching of Continuous Driving Scenes with Growing Architecture

专知会员服务

11+阅读 · 2022年3月12日

【ICLR2022】时序对齐预测的监督表示学习与少样本序列分类

【ICLR2022】时序对齐预测的监督表示学习与少样本序列分类

专知会员服务

21+阅读 · 2022年2月5日

FlowQA: Grasping Flow in History for Conversational Machine Comprehension

FlowQA: Grasping Flow in History for Conversational Machine Comprehension

专知会员服务

34+阅读 · 2019年10月18日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

【NIPS2019】Infidelity and Sensitivity：模型可解释性方法的定量评估

【NIPS2019】Infidelity and Sensitivity：模型可解释性方法的定量评估

AINLP

19+阅读 · 2020年6月14日

【复旦大学】最新《预训练语言模型》2020综述论文大全，50+PTMs分类体系，25页pdf205篇参考文献

【复旦大学】最新《预训练语言模型》2020综述论文大全，50+PTMs分类体系，25页pdf205篇参考文献

专知

22+阅读 · 2020年3月19日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

每日论文 | 用于紧凑语义分割模型的框架搜索；用深度学习进行命名实体消歧；多特征文本风格迁移

每日论文 | 用于紧凑语义分割模型的框架搜索；用深度学习进行命名实体消歧；多特征文本风格迁移

论智

11+阅读 · 2018年11月5日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

《pyramid Attention Network for Semantic Segmentation》

《pyramid Attention Network for Semantic Segmentation》

统计学习与视觉计算组

44+阅读 · 2018年8月30日

用 LDA 和 LSA 两种方法来降维和做 Topic 建模

用 LDA 和 LSA 两种方法来降维和做 Topic 建模

AI研习社

13+阅读 · 2018年8月24日

STRCF for Visual Object Tracking

STRCF for Visual Object Tracking

统计学习与视觉计算组

15+阅读 · 2018年5月29日

Focal Loss for Dense Object Detection

Focal Loss for Dense Object Detection

统计学习与视觉计算组

12+阅读 · 2018年3月15日

From Softmax to Sparsemax-ICML16（1）

From Softmax to Sparsemax-ICML16（1）

KingsGarden

74+阅读 · 2016年11月26日

RC框架-框桁式复合墙混合抗侧力体系抗震性能研究

国家自然科学基金

0+阅读 · 2015年12月31日

基于反馈型级联连接模型的多模态语义SFM方法研究

国家自然科学基金

2+阅读 · 2015年12月31日

钢筋混凝土复杂受力构件的配筋设计方法研究

国家自然科学基金

0+阅读 · 2015年12月31日

考虑材料分布不确定性的结构拓扑优化问题数学建模与求解方法

国家自然科学基金

0+阅读 · 2015年12月31日

基于自适应采样和变复杂度近似的多学科稳健性设计优化方法研究

国家自然科学基金

0+阅读 · 2015年12月31日

四步法三维编织复合材料弯曲疲劳失效多尺度损伤模型

国家自然科学基金

0+阅读 · 2015年12月31日

基于多层次结构特征的新型混杂纤维增强水泥基复合材料(HyFRCC)的性能及机理研究

国家自然科学基金

0+阅读 · 2014年12月31日

一种全新的结构修改重分析方法及其应用

国家自然科学基金

0+阅读 · 2014年12月31日

基于温度效应的CFRP加固钢结构疲劳性能与控制方法研究

国家自然科学基金

0+阅读 · 2014年12月31日

预制装配型钢混凝土梁受力行为与设计方法研究

国家自然科学基金

0+阅读 · 2014年12月31日

Scaling Laws for Moral Machine Judgment in Large Language Models

Arxiv

0+阅读 · 4月30日

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training

Arxiv

0+阅读 · 4月29日

Language-Integrated Recursive Queries (Full Version)

Arxiv

0+阅读 · 4月24日

Pause or Fabricate? Training Language Models for Grounded Reasoning

Arxiv

0+阅读 · 4月21日

An Algorithm-to-Contract Framework without Demand Queries

Arxiv

0+阅读 · 4月9日

Prosocial Persuasion at Scale? Large Language Models Outperform Humans in Donation Appeals Across Levels of Personalization

Arxiv

0+阅读 · 4月3日

Rethinking Language Model Scaling under Transferable Hypersphere Optimization

Arxiv

0+阅读 · 3月30日

On-Policy Context Distillation for Language Models

Arxiv

0+阅读 · 3月23日

A Theory of Composable Lingos for Protocol Dialects

Arxiv

0+阅读 · 3月20日

Leveraging Large Language Models for Generalizing Peephole Optimizations

Arxiv

0+阅读 · 3月19日

VIP会员

文章信息

相关主题

最新内容

DeepSeek 版Claude Code，免费小白安装教程来了！

DeepSeek 版Claude Code，免费小白安装教程来了！

专知会员服务

6+阅读 · 5月5日

【ICML Spotlight 2026】 T²PO: 不确定性引导的探索控制框架，实现稳定多轮Agentic强化学习

【ICML Spotlight 2026】 T²PO: 不确定性引导的探索控制框架，实现稳定多轮Agentic强化学习

专知会员服务

2+阅读 · 5月5日

基础模型驱动的工业智能体：技术成熟度、能力变迁与未竟之挑战

基础模型驱动的工业智能体：技术成熟度、能力变迁与未竟之挑战

专知会员服务

0+阅读 · 5月5日

《机动炮兵的演进与未来：技术进步、历史沿革与炮兵作战前瞻》

《机动炮兵的演进与未来：技术进步、历史沿革与炮兵作战前瞻》

专知会员服务

3+阅读 · 5月5日

《火炮弹药快速效能建模：提升互操作性与技术优势》（报告）

《火炮弹药快速效能建模：提升互操作性与技术优势》（报告）

专知会员服务

4+阅读 · 5月5日

《美空军条令出版物 2-0：情报（2026版）》

《美空军条令出版物 2-0：情报（2026版）》

专知会员服务

9+阅读 · 5月5日

美陆军“飞蝇陷阱5.0”项目将新兴技术交到作战人员手中

美陆军“飞蝇陷阱5.0”项目将新兴技术交到作战人员手中

专知会员服务

3+阅读 · 5月5日

帕兰提尔 Gotham：一个游戏规则改变器

帕兰提尔 Gotham：一个游戏规则改变器

专知会员服务

5+阅读 · 5月5日

【ICML 2026】用测试时训练线性化视觉Transformer：T⁵ 实现 Softmax 注意力到线性复杂度的快速转换

【ICML 2026】用测试时训练线性化视觉Transformer：T⁵ 实现 Softmax 注意力到线性复杂度的快速转换

专知会员服务

2+阅读 · 5月5日

【AAAI 2026】大模型做知识蒸馏：CMM将LLM特征拆解给小模型协同学习

【AAAI 2026】大模型做知识蒸馏：CMM将LLM特征拆解给小模型协同学习

专知会员服务

2+阅读 · 5月5日

【ICML Spotlight 2026 】NonZero：交互引导探索的多智能体蒙特卡洛树搜索

【ICML Spotlight 2026 】NonZero：交互引导探索的多智能体蒙特卡洛树搜索

专知会员服务

8+阅读 · 5月4日

【综述】机器人学习中的世界模型：全面综述

【综述】机器人学习中的世界模型：全面综述

专知会员服务

10+阅读 · 5月4日

伊朗的导弹-无人机行动及其对美国威慑的影响

伊朗的导弹-无人机行动及其对美国威慑的影响

专知会员服务

8+阅读 · 5月4日

《未来战术无人机系统案例研究：量身定制采办策略方法》100页报告

《未来战术无人机系统案例研究：量身定制采办策略方法》100页报告

专知会员服务

8+阅读 · 5月4日

战争贩子：2026年第一季度美国对中东潜在军售激增

战争贩子：2026年第一季度美国对中东潜在军售激增

专知会员服务

6+阅读 · 5月4日

相关VIP内容

不可错过！EPFL《训练大语言模型》课程

不可错过！EPFL《训练大语言模型》课程

专知会员服务

18+阅读 · 2025年4月25日

最高9.0分！这16篇最高分ICLR2025论文必看！从生成模型到MOE等

最高9.0分！这16篇最高分ICLR2025论文必看！从生成模型到MOE等

专知会员服务

26+阅读 · 2024年11月19日

视频大模型中视觉上下文表示的scaling law

视频大模型中视觉上下文表示的scaling law

专知会员服务

24+阅读 · 2024年10月21日

【TPAMI2024】结构化时空对齐视频-语言表示

【TPAMI2024】结构化时空对齐视频-语言表示

专知会员服务

29+阅读 · 2024年10月20日

ACL2024 | IEPILE:大规模基于Schema的信息抽取语料库

ACL2024 | IEPILE:大规模基于Schema的信息抽取语料库

专知会员服务

32+阅读 · 2024年6月20日

国家标准《人工智能预训练模型第3 部分服务能力成熟度评估》

国家标准《人工智能预训练模型第3 部分服务能力成熟度评估》

专知会员服务

63+阅读 · 2024年6月16日

【CVPR 2022】连续驾驶场景与不断增长的建筑的连续立体匹配，Continual Stereo Matching of Continuous Driving Scenes with Growing Architecture

【CVPR 2022】连续驾驶场景与不断增长的建筑的连续立体匹配，Continual Stereo Matching of Continuous Driving Scenes with Growing Architecture

专知会员服务

11+阅读 · 2022年3月12日

【ICLR2022】时序对齐预测的监督表示学习与少样本序列分类

【ICLR2022】时序对齐预测的监督表示学习与少样本序列分类

专知会员服务

21+阅读 · 2022年2月5日

FlowQA: Grasping Flow in History for Conversational Machine Comprehension

FlowQA: Grasping Flow in History for Conversational Machine Comprehension

专知会员服务

34+阅读 · 2019年10月18日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

热门VIP内容

开通专知VIP会员享更多权益服务

【ICML Spotlight 2026】 T²PO: 不确定性引导的探索控制框架，实现稳定多轮Agentic强化学习

《机动炮兵的演进与未来：技术进步、历史沿革与炮兵作战前瞻》

DeepSeek 版Claude Code，免费小白安装教程来了！

基础模型驱动的工业智能体：技术成熟度、能力变迁与未竟之挑战

相关资讯

【NIPS2019】Infidelity and Sensitivity：模型可解释性方法的定量评估

【NIPS2019】Infidelity and Sensitivity：模型可解释性方法的定量评估

AINLP

19+阅读 · 2020年6月14日

【复旦大学】最新《预训练语言模型》2020综述论文大全，50+PTMs分类体系，25页pdf205篇参考文献

【复旦大学】最新《预训练语言模型》2020综述论文大全，50+PTMs分类体系，25页pdf205篇参考文献

专知

22+阅读 · 2020年3月19日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

每日论文 | 用于紧凑语义分割模型的框架搜索；用深度学习进行命名实体消歧；多特征文本风格迁移

每日论文 | 用于紧凑语义分割模型的框架搜索；用深度学习进行命名实体消歧；多特征文本风格迁移

论智

11+阅读 · 2018年11月5日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

《pyramid Attention Network for Semantic Segmentation》

《pyramid Attention Network for Semantic Segmentation》

统计学习与视觉计算组

44+阅读 · 2018年8月30日

用 LDA 和 LSA 两种方法来降维和做 Topic 建模

用 LDA 和 LSA 两种方法来降维和做 Topic 建模

AI研习社

13+阅读 · 2018年8月24日

STRCF for Visual Object Tracking

STRCF for Visual Object Tracking

统计学习与视觉计算组

15+阅读 · 2018年5月29日

Focal Loss for Dense Object Detection

Focal Loss for Dense Object Detection

统计学习与视觉计算组

12+阅读 · 2018年3月15日

From Softmax to Sparsemax-ICML16（1）

From Softmax to Sparsemax-ICML16（1）

KingsGarden

74+阅读 · 2016年11月26日

相关论文

Scaling Laws for Moral Machine Judgment in Large Language Models

Arxiv

0+阅读 · 4月30日

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training

Arxiv

0+阅读 · 4月29日

Language-Integrated Recursive Queries (Full Version)

Arxiv

0+阅读 · 4月24日

Pause or Fabricate? Training Language Models for Grounded Reasoning

Arxiv

0+阅读 · 4月21日

An Algorithm-to-Contract Framework without Demand Queries

Arxiv

0+阅读 · 4月9日

Prosocial Persuasion at Scale? Large Language Models Outperform Humans in Donation Appeals Across Levels of Personalization

Arxiv

0+阅读 · 4月3日

Rethinking Language Model Scaling under Transferable Hypersphere Optimization

Arxiv

0+阅读 · 3月30日

On-Policy Context Distillation for Language Models

Arxiv

0+阅读 · 3月23日

A Theory of Composable Lingos for Protocol Dialects

Arxiv

0+阅读 · 3月20日

Leveraging Large Language Models for Generalizing Peephole Optimizations

Arxiv

0+阅读 · 3月19日

相关基金

RC框架-框桁式复合墙混合抗侧力体系抗震性能研究

国家自然科学基金

0+阅读 · 2015年12月31日

基于反馈型级联连接模型的多模态语义SFM方法研究

国家自然科学基金

2+阅读 · 2015年12月31日

钢筋混凝土复杂受力构件的配筋设计方法研究

国家自然科学基金

0+阅读 · 2015年12月31日

考虑材料分布不确定性的结构拓扑优化问题数学建模与求解方法

国家自然科学基金

0+阅读 · 2015年12月31日

基于自适应采样和变复杂度近似的多学科稳健性设计优化方法研究

国家自然科学基金

0+阅读 · 2015年12月31日

四步法三维编织复合材料弯曲疲劳失效多尺度损伤模型

国家自然科学基金

0+阅读 · 2015年12月31日

基于多层次结构特征的新型混杂纤维增强水泥基复合材料(HyFRCC)的性能及机理研究

国家自然科学基金

0+阅读 · 2014年12月31日

一种全新的结构修改重分析方法及其应用

国家自然科学基金

0+阅读 · 2014年12月31日

基于温度效应的CFRP加固钢结构疲劳性能与控制方法研究

国家自然科学基金

0+阅读 · 2014年12月31日

预制装配型钢混凝土梁受力行为与设计方法研究

国家自然科学基金

0+阅读 · 2014年12月31日

微信扫码咨询专知VIP会员