We study Leaky ResNets, which interpolate between ResNets and Fully-Connected nets depending on an 'effective depth' hyper-parameter $\tilde{L}$. In the infinite depth limit, we study 'representation geodesics' $A_{p}$: continuous paths in representation space (similar to NeuralODEs) from input $p=0$ to output $p=1$ that minimize the parameter norm of the network. We give a Lagrangian and Hamiltonian reformulation, which highlight the importance of two terms: a kinetic energy which favors small layer derivatives $\partial_{p}A_{p}$ and a potential energy that favors low-dimensional representations, as measured by the 'Cost of Identity'. The balance between these two forces offers an intuitive understanding of feature learning in ResNets. We leverage this intuition to explain the emergence of a bottleneck structure, as observed in previous work: for large $\tilde{L}$ the potential energy dominates and leads to a separation of timescales, where the representation jumps rapidly from the high dimensional inputs to a low-dimensional representation, move slowly inside the space of low-dimensional representations, before jumping back to the potentially high-dimensional outputs. Inspired by this phenomenon, we train with an adaptive layer step-size to adapt to the separation of timescales.


翻译:我们研究泄漏残差网络(Leaky ResNets),该网络通过'有效深度'超参数$\tilde{L}$在残差网络(ResNets)与全连接网络之间插值。在无限深度极限下,我们研究'表示测地线'$A_{p}$:表示空间中从输入$p=0$到输出$p=1$的连续路径(类似于神经常微分方程),该路径最小化网络的参数范数。我们提出拉格朗日与哈密顿重构,突显两项关键作用:倾向于小层导数$\partial_{p}A_{p}$的动能,以及通过'恒等代价'衡量、倾向于低维表示势能。这两力平衡为残差网络中的特征学习提供了直观理解。我们利用该直觉解释先前研究中观察到的瓶颈结构涌现现象:当$\tilde{L}$较大时,势能主导并导致时间尺度分离——表示从高维输入快速跃迁至低维表示,在低维表示空间缓慢移动,再跳回可能的高维输出。受此现象启发,我们采用自适应层步长训练以适应时间尺度分离。

0
下载
关闭预览

相关内容

【牛津大学博士论文】深度学习算法的渐近分析,186页pdf
专知会员服务
20+阅读 · 2021年5月30日
神经网络的拓扑结构,TOPOLOGY OF DEEP NEURAL NETWORKS
专知会员服务
35+阅读 · 2020年4月15日
连载▍AlexNet结构详解(引用MrGiovanni博士)
36大数据
10+阅读 · 2019年3月28日
从信息瓶颈理论一瞥机器学习的“大一统理论”
网络表示学习介绍
人工智能前沿讲习班
18+阅读 · 2018年11月26日
手把手教你构建ResNet残差网络
专知
38+阅读 · 2018年4月27日
【干货】Lossless Triplet Loss: 一种高效的Siamese网络损失函数
机器学习研究会
29+阅读 · 2018年2月21日
图上的归纳表示学习
科技创新与创业
23+阅读 · 2017年11月9日
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
VIP会员
最新内容
《决策模型比较研究》
专知会员服务
7+阅读 · 今天5:16
《美军水下战与海床战概述及本地实施》
专知会员服务
3+阅读 · 今天4:30
面向未来冲突推进陆军情报体制改革
专知会员服务
3+阅读 · 今天4:12
乌克兰纵深打击如何重塑俄罗斯的战略选择
专知会员服务
3+阅读 · 7月24日
俄乌战争中关于中程打击无人机部署的经验启示
《基于强化学习的自动化红队测试》
专知会员服务
5+阅读 · 7月23日
相关VIP内容
【牛津大学博士论文】深度学习算法的渐近分析,186页pdf
专知会员服务
20+阅读 · 2021年5月30日
神经网络的拓扑结构,TOPOLOGY OF DEEP NEURAL NETWORKS
专知会员服务
35+阅读 · 2020年4月15日
相关资讯
连载▍AlexNet结构详解(引用MrGiovanni博士)
36大数据
10+阅读 · 2019年3月28日
从信息瓶颈理论一瞥机器学习的“大一统理论”
网络表示学习介绍
人工智能前沿讲习班
18+阅读 · 2018年11月26日
手把手教你构建ResNet残差网络
专知
38+阅读 · 2018年4月27日
【干货】Lossless Triplet Loss: 一种高效的Siamese网络损失函数
机器学习研究会
29+阅读 · 2018年2月21日
图上的归纳表示学习
科技创新与创业
23+阅读 · 2017年11月9日
相关基金
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
Top
微信扫码咨询专知VIP会员