In this work, we present an integrated system for spatiotemporal summarization of 360-degrees videos. The video summary production mainly involves the detection of salient events and their synopsis into a concise summary. The analysis relies on state-of-the-art methods for saliency detection in 360-degrees video (ATSal and SST-Sal) and video summarization (CA-SUM). It also contains a mechanism that classifies a 360-degrees video based on the use of static or moving camera during recording and decides which saliency detection method will be used, as well as a 2D video production component that is responsible to create a conventional 2D video containing the salient events in the 360-degrees video. Quantitative evaluations using two datasets for 360-degrees video saliency detection (VR-EyeTracking, Sports-360) show the accuracy and positive impact of the developed decision mechanism, and justify our choice to use two different methods for detecting the salient events. A qualitative analysis using content from these datasets, gives further insights about the functionality of the decision mechanism, shows the pros and cons of each used saliency detection method and demonstrates the advanced performance of the trained summarization method against a more conventional approach.
翻译:本文提出了一种面向360度视频时空摘要的集成系统。视频摘要生成主要涉及显著性事件检测及其浓缩摘要合成。该分析基于360度视频显著性检测(ATSal和SST-Sal)及视频摘要(CA-SUM)领域的先进方法。系统还包含一个分类机制,用于根据录制时静态或动态摄像机的使用情况对360度视频进行分类,并决定采用哪种显著性检测方法;同时配备2D视频生成组件,负责将360度视频中的显著性事件转化为传统2D视频。基于两个360度视频显著性检测数据集(VR-EyeTracking、Sports-360)的定量评估表明,所开发的决策机制具有准确性与积极效果,验证了采用两种不同方法检测显著性事件的合理性。基于这些数据集内容的定性分析进一步揭示了决策机制的运行机理,展示了每种显著性检测方法的优缺点,并证明了经过训练的摘要方法相较于传统方法的优越性能。