Existing single-modal and multi-modal salient object detection (SOD) methods focus on designing specific architectures tailored for their respective tasks. However, developing completely different models for different tasks leads to labor and time consumption, as well as high computational and practical deployment costs. In this paper, we make the first attempt to address both single-modal and multi-modal SOD in a unified framework called UniSOD. Nevertheless, assigning appropriate strategies to modality variable inputs is challenging. To this end, UniSOD learns modality-aware prompts with task-specific hints through adaptive prompt learning, which are plugged into the proposed pre-trained baseline SOD model to handle corresponding tasks, while only requiring few learnable parameters compared to training the entire model. Each modality-aware prompt is generated from a switchable prompt generation block, which performs structural switching solely relied on single-modal and multi-modal inputs. UniSOD achieves consistent performance improvement on 14 benchmark datasets for RGB, RGB-D, and RGB-T SOD, which demonstrates that our method effectively and efficiently unifies single-modal and multi-modal SOD tasks.
翻译:现有的单模态与多模态显著性目标检测方法专注于针对各自任务设计特定架构。然而,为不同任务开发完全不同的模型会导致人力与时间消耗,并带来高昂的计算和实际部署成本。本文首次尝试在统一框架UniSOD中同时解决单模态与多模态显著性目标检测问题。尽管如此,为模态变量输入分配适当策略仍具有挑战性。为此,UniSOD通过自适应提示学习学习带有任务特定线索的模态感知提示,并将其嵌入所提出的预训练基线显著性目标检测模型以处理相应任务,相较于训练完整模型仅需少量可学习参数。每个模态感知提示由可切换提示生成模块生成,该模块仅根据单模态与多模态输入进行结构切换。UniSOD在RGB、RGB-D和RGB-T显著性目标检测的14个基准数据集上均实现了一致的性能提升,表明我们的方法能够有效且高效地统一单模态与多模态显著性目标检测任务。