Leveraging the rich information extracted from light field (LF) cameras is instrumental for dense prediction tasks. However, adapting light field data to enhance Salient Object Detection (SOD) still follows the traditional RGB methods and remains under-explored in the community. Previous approaches predominantly employ a custom two-stream design to discover the implicit angular feature within light field cameras, leading to significant information isolation between different LF representations. In this study, we propose an efficient paradigm (LF Tracy) to address this limitation. We eschew the conventional specialized fusion and decoder architecture for a dual-stream backbone in favor of a unified, single-pipeline approach. This comprises firstly a simple yet effective data augmentation strategy called MixLD to bridge the connection of spatial, depth, and implicit angular information under different LF representations. A highly efficient information aggregation (IA) module is then introduced to boost asymmetric feature-wise information fusion. Owing to this innovative approach, our model surpasses the existing state-of-the-art methods, particularly demonstrating a 23% improvement over previous results on the latest large-scale PKU dataset. By utilizing only 28.9M parameters, the model achieves a 10% increase in accuracy with 3M additional parameters compared to its backbone using RGB images and an 86% rise to its backbone using LF images. The source code will be made publicly available at https://github.com/FeiBryantkit/LF-Tracy.
翻译:摘要:利用光场相机提取的丰富信息对密集预测任务至关重要。然而,将光场数据适配以增强显著目标检测仍遵循传统RGB方法,且在学界探索不足。现有方法大多采用定制的双流架构以挖掘光场相机中的隐式角度特征,导致不同光场表征之间存在显著的信息隔离。本研究提出一种高效范式(LF Tracy)以解决此限制。我们摒弃了传统双流主干网络专用的特征融合与解码器架构,转而采用统一的单流水线方法。该方法首先包含一种简单高效的数据增强策略MixLD,用于桥接不同光场表征下空间、深度及隐式角度信息之间的关联;随后引入高效信息聚合模块以增强非对称特征层面的信息融合。凭借这一创新方法,我们的模型超越了现有最先进技术,特别是在最新大规模PKU数据集上相较于先前结果实现了23%的提升。该模型仅使用2890万参数,相比其基于RGB图像的主干网络,仅增加300万参数即可实现10%的精度提升;而相比其基于光场图像的主干网络,精度提升达86%。源代码将公开于https://github.com/FeiBryantkit/LF-Tracy。