Implicit deep learning has recently gained popularity with applications ranging from meta-learning to Deep Equilibrium Networks (DEQs). In its general formulation, it relies on expressing some components of deep learning pipelines implicitly, typically via a root equation called the inner problem. In practice, the solution of the inner problem is approximated during training with an iterative procedure, usually with a fixed number of inner iterations. During inference, the inner problem needs to be solved with new data. A popular belief is that increasing the number of inner iterations compared to the one used during training yields better performance. In this paper, we question such an assumption and provide a detailed theoretical analysis in a simple setting. We demonstrate that overparametrization plays a key role: increasing the number of iterations at test time cannot improve performance for overparametrized networks. We validate our theory on an array of implicit deep-learning problems. DEQs, which are typically overparametrized, do not benefit from increasing the number of iterations at inference while meta-learning, which is typically not overparametrized, benefits from it.
翻译:隐式深度学习近期在元学习到深度平衡网络(DEQs)等应用中备受关注。其一般形式依赖于将深度学习管道的某些组件隐式表达,通常通过称为内问题的根方程实现。实践中,内问题的解在训练期间通过迭代过程近似求解,通常采用固定数量的内迭代次数。推理时,需用新数据求解内问题。一种普遍观点认为,相比训练时使用的迭代次数,增加内迭代次数可带来更优性能。本文质疑这一假设,并在简单场景中提供详细的理论分析。我们证明过参数化起关键作用:对于过参数化网络,在测试时增加迭代次数无法提升性能。我们在系列隐式深度学习问题上验证了该理论。典型的过参数化模型DEQs在推理时增加迭代次数无益,而通常非过参数化的元学习则能从中获益。