Representing human performance at high-fidelity is an essential building block in diverse applications, such as film production, computer games or videoconferencing. To close the gap to production-level quality, we introduce HumanRF, a 4D dynamic neural scene representation that captures full-body appearance in motion from multi-view video input, and enables playback from novel, unseen viewpoints. Our novel representation acts as a dynamic video encoding that captures fine details at high compression rates by factorizing space-time into a temporal matrix-vector decomposition. This allows us to obtain temporally coherent reconstructions of human actors for long sequences, while representing high-resolution details even in the context of challenging motion. While most research focuses on synthesizing at resolutions of 4MP or lower, we address the challenge of operating at 12MP. To this end, we introduce ActorsHQ, a novel multi-view dataset that provides 12MP footage from 160 cameras for 16 sequences with high-fidelity, per-frame mesh reconstructions. We demonstrate challenges that emerge from using such high-resolution data and show that our newly introduced HumanRF effectively leverages this data, making a significant step towards production-level quality novel view synthesis.
翻译:以高保真度表现人类动态表现是电影制作、电子游戏或视频会议等多样化应用中的关键组成部分。为缩小与制作级质量的差距,我们提出HumanRF——一种能够从多视角视频输入中捕捉全身运动外观,并支持从全新未见视角回放的4D动态神经场景表示。该新表示方法通过将时空分解为时间矩阵-向量分解,以高压缩率编码精细细节,实现动态视频编码。这使得我们能够对长序列中的人类演员进行时间一致的动态重建,即使在复杂运动背景下也能呈现高分辨率细节。尽管多数研究聚焦于4MP或更低分辨率合成,我们挑战了12MP分辨率下的操作难题。为此,我们提出ActorsHQ——一个包含16个序列、由160台摄像机拍摄的12MP视频素材及每帧高保真网格重建的新型多视角数据集。我们展示了使用此类高分辨率数据时涌现的挑战,并证明新提出的HumanRF能有效利用该数据,为迈向制作级质量的新视角合成迈出重要一步。