Radar-camera BEV perception often suffers from degraded performance when evaluated across datasets, as changes in driving scenes, sensor configurations, and environmental conditions can alter both the input observations and the internal fused representations. This work studies this issue from the perspective of source-domain variation modeling, aiming to improve the robustness of BEV-based 3D detectors without relying on target-domain samples. We introduce a framework that characterizes visual scene variations in the frequency domain and uses them to synthesize diverse source-domain views. By comparing the resulting fused BEV representations, the framework further captures how image-level variations influence multi-modal BEV features. These variation patterns are then used to regularize the detector, encouraging the learned fusion space to remain stable under latent scene changes. The proposed method is applied only during training and leaves the inference pipeline unchanged. Experiments on cross-dataset radar-camera 3D detection between View-of-Delft and TJ4DRadSet demonstrate consistent improvements over multiple BEV fusion backbones, and the gains remain effective when a small amount of target-domain data is available.
翻译:雷达-相机鸟瞰图感知在跨数据集评估时常常性能下降,因为驾驶场景、传感器配置和环境条件的变化会同时改变输入观测和内部融合表示。本研究从源域变化建模的角度探讨该问题,旨在提升基于BEV的三维检测器的鲁棒性,且无需依赖目标域样本。我们提出一个框架,在频域中表征视觉场景变化,并利用这些变化合成多样化的源域视图。通过比较生成的融合BEV表示,该框架进一步捕捉图像级变化如何影响多模态BEV特征。这些变化模式随后被用于正则化检测器,促使学习到的融合空间在潜在场景变化下保持稳定。所提方法仅在训练阶段使用,推理流程保持不变。在View-of-Delft与TJ4DRadSet之间的跨数据集雷达-相机三维检测实验表明,该方法在多个BEV融合骨干网络上均能实现一致提升,且当少量目标域数据可用时,性能增益依然有效。