We propose a novel generative saliency prediction framework that adopts an informative energy-based model as a prior distribution. The energy-based prior model is defined on the latent space of a saliency generator network that generates the saliency map based on a continuous latent variables and an observed image. Both the parameters of saliency generator and the energy-based prior are jointly trained via Markov chain Monte Carlo-based maximum likelihood estimation, in which the sampling from the intractable posterior and prior distributions of the latent variables are performed by Langevin dynamics. With the generative saliency model, we can obtain a pixel-wise uncertainty map from an image, indicating model confidence in the saliency prediction. Different from existing generative models, which define the prior distribution of the latent variables as a simple isotropic Gaussian distribution, our model uses an energy-based informative prior which can be more expressive in capturing the latent space of the data. With the informative energy-based prior, we extend the Gaussian distribution assumption of generative models to achieve a more representative distribution of the latent space, leading to more reliable uncertainty estimation. We apply the proposed frameworks to both RGB and RGB-D salient object detection tasks with both transformer and convolutional neural network backbones. We further propose an adversarial learning algorithm and a variational inference algorithm as alternatives to train the proposed generative framework. Experimental results show that our generative saliency model with an energy-based prior can achieve not only accurate saliency predictions but also reliable uncertainty maps that are consistent with human perception. Results and code are available at \url{https://github.com/JingZhang617/EBMGSOD}.
翻译:我们提出了一种新颖的生成式显著性预测框架,该框架采用基于能量的信息型先验模型作为先验分布。该能量先验模型定义在显著性生成网络的潜变量空间上,该网络基于连续潜变量和观测图像生成显著性图。通过马尔可夫链蒙特卡洛最大似然估计联合训练显著性生成器的参数与能量先验,其中对潜变量难以处理的先验和后验分布采用朗之万动力学进行采样。借助该生成式显著性模型,我们可从图像中获取像素级的不确定性图,反映模型在显著性预测中的置信度。与将潜变量先验分布定义为简单各向同性高斯分布的现有生成模型不同,我们的模型采用更具表达能力的能量信息型先验,能更有效地捕捉数据的潜变量空间。通过引入信息型能量先验,我们扩展了生成模型中的高斯分布假设,使潜变量分布更具代表性,从而实现更可靠的不确定性估计。我们将所提框架同时应用于基于Transformer和卷积神经网络(CNN)骨干网络的RGB与RGB-D显著目标检测任务。此外,我们提出对抗学习算法和变分推理算法作为训练该生成框架的替代方案。实验结果表明,基于能量先验的生成式显著性模型不仅能实现精准的显著性预测,还能生成与人类感知一致的可信不确定性图。相关结果与代码已开源至 \url{https://github.com/JingZhang617/EBMGSOD}。