Understanding semantic scene segmentation of urban scenes captured from the Unmanned Aerial Vehicles (UAV) perspective plays a vital role in building a perception model for UAV. With the limitations of large-scale densely labeled data, semantic scene segmentation for UAV views requires a broad understanding of an object from both its top and side views. Adapting from well-annotated autonomous driving data to unlabeled UAV data is challenging due to the cross-view differences between the two data types. Our work proposes a novel Cross-View Adaptation (CROVIA) approach to effectively adapt the knowledge learned from on-road vehicle views to UAV views. First, a novel geometry-based constraint to cross-view adaptation is introduced based on the geometry correlation between views. Second, cross-view correlations from image space are effectively transferred to segmentation space without any requirement of paired on-road and UAV view data via a new Geometry-Constraint Cross-View (GeiCo) loss. Third, the multi-modal bijective networks are introduced to enforce the global structural modeling across views. Experimental results on new cross-view adaptation benchmarks introduced in this work, i.e., SYNTHIA to UAVID and GTA5 to UAVID, show the State-of-the-Art (SOTA) performance of our approach over prior adaptation methods
翻译:理解从无人机视角捕获的城市场景的语义分割在构建无人机感知模型中起着关键作用。由于大规模密集标注数据的限制,无人机视角的语义场景分割需要从俯视图和侧视图两方面对物体有广泛的理解。从标注良好的自动驾驶数据适应到未标注的无人机数据因两种数据类型间的跨视角差异而具有挑战性。我们的工作提出了一种新颖的跨视角适应方法,以有效将从道路车辆视角学习到的知识迁移到无人机视角。首先,基于视角间的几何相关性,引入了一种新颖的几何约束用于跨视角适应。其次,通过一种新的几何约束跨视角损失,将图像空间中的跨视角相关性有效迁移到分割空间,无需成对的道路和无人机视角数据。第三,引入多模态双射网络以强化跨视角的全局结构建模。在本文提出的新型跨视角适应基准(即SYNTHIA到UAVID和GTA5到UAVID)上的实验结果表明,我们的方法相较于先前的适应方法达到了最先进的性能。