Learning-based visual odometry (VO) algorithms achieve remarkable performance on common static scenes, benefiting from high-capacity models and massive annotated data, but tend to fail in dynamic, populated environments. Semantic segmentation is largely used to discard dynamic associations before estimating camera motions but at the cost of discarding static features and is hard to scale up to unseen categories. In this paper, we leverage the mutual dependence between camera ego-motion and motion segmentation and show that both can be jointly refined in a single learning-based framework. In particular, we present DytanVO, the first supervised learning-based VO method that deals with dynamic environments. It takes two consecutive monocular frames in real-time and predicts camera ego-motion in an iterative fashion. Our method achieves an average improvement of 27.7% in ATE over state-of-the-art VO solutions in real-world dynamic environments, and even performs competitively among dynamic visual SLAM systems which optimize the trajectory on the backend. Experiments on plentiful unseen environments also demonstrate our method's generalizability.
翻译:基于学习的视觉里程计算法凭借高容量模型和大量标注数据在常见静态场景中表现出色,但在动态、拥挤环境中容易失效。语义分割通常被用于丢弃动态关联信息以估计相机运动,但这会导致静态特征丢失且难以扩展到未见过的类别。本文利用相机自运动与运动分割之间的相互依赖性,证明两者可以在单一的学习框架中联合优化。具体而言,我们提出DytanVO——首个处理动态环境的监督式视觉里程计方法。该方法实时处理连续两帧单目图像,以迭代方式预测相机自运动。在真实动态环境中,我们的方法相较于最新视觉里程计算法的平均绝对轨迹误差降低了27.7%,甚至与后端优化轨迹的动态视觉SLAM系统相比也具备竞争力。在大量未见环境上的实验进一步验证了该方法的泛化能力。