Traditional visual-inertial state estimation targets absolute camera poses and spatial landmark locations while first-order kinematics are typically resolved as an implicitly estimated sub-state. However, this poses a risk in velocity-based control scenarios, as the quality of the estimation of kinematics depends on the stability of absolute camera and landmark coordinates estimation. To address this issue, we propose a novel solution to tight visual-inertial fusion directly at the level of first-order kinematics by employing a dynamic vision sensor instead of a normal camera. More specifically, we leverage trifocal tensor geometry to establish an incidence relation that directly depends on events and camera velocity, and demonstrate how velocity estimates in highly dynamic situations can be obtained over short time intervals. Noise and outliers are dealt with using a nested two-layer RANSAC scheme. Additionally, smooth velocity signals are obtained from a tight fusion with pre-integrated inertial signals using a sliding window optimizer. Experiments on both simulated and real data demonstrate that the proposed tight event-inertial fusion leads to continuous and reliable velocity estimation in highly dynamic scenarios independently of absolute coordinates. Furthermore, in extreme cases, it achieves more stable and more accurate estimation of kinematics than traditional, point-position-based visual-inertial odometry.
翻译:传统视觉-惯性状态估计以绝对相机位姿和空间路标点位置为目标,而一阶运动学量通常作为隐式估计的子状态处理。然而,在基于速度的控制场景中,这种做法存在风险,因为运动学估计质量依赖于绝对相机坐标和路标坐标估计的稳定性。为解决这一问题,我们提出一种新颖的视觉-惯性紧融合方案,直接在一阶运动学层面实现融合,采用动态视觉传感器替代传统相机。具体而言,我们利用三焦张量几何建立直接依赖事件与相机速度的关联关系,并演示如何在短时间内获取高动态场景下的速度估计。噪声与离群点通过嵌套双层RANSAC方案处理。此外,通过滑动窗口优化器与预积分惯性信号的紧融合,可获得平滑的速度信号。在仿真与真实数据上的实验表明,所提事件-惯性紧融合方法能在高动态场景中独立于绝对坐标实现连续可靠的速度估计。在极端情况下,相比基于点位置的视觉-惯性里程计,该方法能获得更稳定、更精确的运动学估计。