The bio-inspired event cameras or dynamic vision sensors are capable of asynchronously capturing per-pixel brightness changes (called event-streams) in high temporal resolution and high dynamic range. However, the non-structural spatial-temporal event-streams make it challenging for providing intuitive visualization with rich semantic information for human vision. It calls for events-to-video (E2V) solutions which take event-streams as input and generate high quality video frames for intuitive visualization. However, current solutions are predominantly data-driven without considering the prior knowledge of the underlying statistics relating event-streams and video frames. It highly relies on the non-linearity and generalization capability of the deep neural networks, thus, is struggling on reconstructing detailed textures when the scenes are complex. In this work, we propose \textbf{E2HQV}, a novel E2V paradigm designed to produce high-quality video frames from events. This approach leverages a model-aided deep learning framework, underpinned by a theory-inspired E2V model, which is meticulously derived from the fundamental imaging principles of event cameras. To deal with the issue of state-reset in the recurrent components of E2HQV, we also design a temporal shift embedding module to further improve the quality of the video frames. Comprehensive evaluations on the real world event camera datasets validate our approach, with E2HQV, notably outperforming state-of-the-art approaches, e.g., surpassing the second best by over 40\% for some evaluation metrics.
翻译:受生物启发的事件相机或动态视觉传感器能够以高时间分辨率和高动态范围异步捕捉每个像素的亮度变化(称为事件流)。然而,非结构化的时空事件流为人类视觉提供具有丰富语义信息的直观可视化带来了挑战。这需要事件到视频(E2V)解决方案,以事件流作为输入,生成高质量的视频帧以实现直观可视化。然而,当前的解决方案主要依赖数据驱动,未考虑事件流与视频帧之间潜在统计关系的先验知识。它们高度依赖于深度神经网络的非线性和泛化能力,因此在场景复杂时难以重建详细纹理。在这项工作中,我们提出了E2HQV,一种旨在从事件生成高质量视频帧的新型E2V范式。该方法利用模型辅助的深度学习框架,该框架以理论启发的E2V模型为基础,该模型从事件相机的基本成像原理中精心推导得出。为解决E2HQV中循环组件的状态重置问题,我们还设计了一个时间平移嵌入模块,以进一步提高视频帧的质量。在真实世界事件相机数据集上的全面评估验证了我们方法的有效性,E2HQV显著优于现有最先进方法,例如,在某些评估指标上超过第二名40%以上。