Deep unrolling, or unfolding, is an emerging learning-to-optimize method that unrolls a truncated iterative algorithm in the layers of a trainable neural network. However, the convergence guarantees and generalizability of the unrolled networks are still open theoretical problems. To tackle these problems, we provide deep unrolled architectures with a stochastic descent nature by imposing descending constraints during training. The descending constraints are forced layer by layer to ensure that each unrolled layer takes, on average, a descent step toward the optimum during training. We theoretically prove that the sequence constructed by the outputs of the unrolled layers is then guaranteed to converge for unseen problems, assuming no distribution shift between training and test problems. We also show that standard unrolling is brittle to perturbations, and our imposed constraints provide the unrolled networks with robustness to additive noise and perturbations. We numerically assess unrolled architectures trained under the proposed constraints in two different applications, including the sparse coding using learnable iterative shrinkage and thresholding algorithm (LISTA) and image inpainting using proximal generative flow (GLOW-Prox), and demonstrate the performance and robustness benefits of the proposed method.
翻译:深度展开(或解折叠)是一种新兴的学习优化方法,它将截断的迭代算法展开到可训练神经网络的各层中。然而,展开网络的收敛保证和泛化能力仍是开放的理论问题。为解决这些问题,我们通过在训练过程中施加下降约束,为深度展开架构赋予了随机下降特性。这些下降约束逐层强制实现,确保每个展开层在训练过程中平均朝向最优解下降一步。我们从理论上证明,由展开层输出构成的序列在未见问题上保证收敛(假设训练问题与测试问题之间无分布偏移)。我们还表明,标准展开对扰动是脆弱的,而所施加的约束为展开网络提供了对加性噪声和扰动的鲁棒性。我们在两种不同应用中数值评估了在提出约束下训练的展开架构,包括使用可学习迭代收缩阈值算法(LISTA)的稀疏编码和使用近端生成流(GLOW-Prox)的图像修复,并展示了所提方法的性能和鲁棒性优势。