The deep unfolding approach has attracted significant attention in computer vision tasks, which well connects conventional image processing modeling manners with more recent deep learning techniques. Specifically, by establishing a direct correspondence between algorithm operators at each implementation step and network modules within each layer, one can rationally construct an almost ``white box'' network architecture with high interpretability. In this architecture, only the predefined component of the proximal operator, known as a proximal network, needs manual configuration, enabling the network to automatically extract intrinsic image priors in a data-driven manner. In current deep unfolding methods, such a proximal network is generally designed as a CNN architecture, whose necessity has been proven by a recent theory. That is, CNN structure substantially delivers the translational invariant image prior, which is the most universally possessed structural prior across various types of images. However, standard CNN-based proximal networks have essential limitations in capturing the rotation symmetry prior, another universal structural prior underlying general images. This leaves a large room for further performance improvement in deep unfolding approaches. To address this issue, this study makes efforts to suggest a high-accuracy rotation equivariant proximal network that effectively embeds rotation symmetry priors into the deep unfolding framework. Especially, we deduce, for the first time, the theoretical equivariant error for such a designed proximal network with arbitrary layers under arbitrary rotation degrees. This analysis should be the most refined theoretical conclusion for such error evaluation to date and is also indispensable for supporting the rationale behind such networks with intrinsic interpretability requirements.
翻译:深度展开方法因其将传统图像处理建模方式与最新深度学习技术有效结合,在计算机视觉任务中备受关注。具体而言,通过建立每个实现步骤中的算法算子与网络各层模块之间的直接对应关系,可以合理构建出具有高可解释性的近乎"白箱"网络架构。在该架构中,仅需手动配置近端算子的预定义组件(即近端网络),使网络能够以数据驱动方式自动提取图像本质先验。当前深度展开方法中,此类近端网络通常被设计为CNN架构,其必要性已得到最新理论验证——CNN结构实质性地提供了平移不变图像先验,这是各类图像中最普遍拥有的结构先验。然而,基于标准CNN的近端网络在捕捉旋转对称先验(通用图像的另一基础结构先验)方面存在本质局限,这为深度展开方法的性能提升留下了较大空间。针对该问题,本研究提出一种高精度旋转等变近端网络,将旋转对称先验有效嵌入深度展开框架。特别地,我们首次推导出该网络在任意层数、任意旋转角度下的理论等变误差,该分析是当前对此类误差评估最精细的理论结论,也为支持具有内在可解释性要求的网络合理性提供不可或缺的理论基础。