Images acquired in hazy conditions have degradations induced in them. Dehazing such images is a vexed and ill-posed problem. Scores of prior-based and learning-based approaches have been proposed to mitigate the effect of haze and generate haze-free images. Many conventional methods are constrained by their lack of awareness regarding scene depth and their incapacity to capture long-range dependencies. In this paper, a method that uses residual learning and vision transformers in an attention module is proposed. It essentially comprises two networks: In the first one, the network takes the ratio of a hazy image and the approximated transmission matrix to estimate a residual map. The second network takes this residual image as input and passes it through convolution layers before superposing it on the generated feature maps. It is then passed through global context and depth-aware transformer encoders to obtain channel attention. The attention module then infers the spatial attention map before generating the final haze-free image. Experimental results, including several quantitative metrics, demonstrate the efficiency and scalability of the suggested methodology.
翻译:雾霾条件下获取的图像存在退化现象。去除此类图像中的雾霾是一个棘手且病态的问题。大量基于先验和基于学习的方法已被提出以减轻雾霾影响并生成无雾图像。许多传统方法受限于缺乏场景深度感知能力以及无法捕捉长程依赖关系。本文提出了一种在注意力模块中结合残差学习与视觉Transformer的方法。该方法主要由两个网络构成:第一个网络利用雾霾图像与近似传输矩阵的比率来估计残差图;第二个网络将该残差图像作为输入,通过卷积层处理后叠加到生成的特征图上,随后通过全局上下文和深度感知Transformer编码器获取通道注意力。注意力模块在生成最终无雾图像前进一步推断空间注意力图。包含多项量化指标在内的实验结果表明,所提方法具有高效性和可扩展性。