Accurately and efficiently modeling dynamic scenes and motions is considered so challenging a task due to temporal dynamics and motion complexity. To address these challenges, we propose DynMF, a compact and efficient representation that decomposes a dynamic scene into a few neural trajectories. We argue that the per-point motions of a dynamic scene can be decomposed into a small set of explicit or learned trajectories. Our carefully designed neural framework consisting of a tiny set of learned basis queried only in time allows for rendering speed similar to 3D Gaussian Splatting, surpassing 120 FPS, while at the same time, requiring only double the storage compared to static scenes. Our neural representation adequately constrains the inherently underconstrained motion field of a dynamic scene leading to effective and fast optimization. This is done by biding each point to motion coefficients that enforce the per-point sharing of basis trajectories. By carefully applying a sparsity loss to the motion coefficients, we are able to disentangle the motions that comprise the scene, independently control them, and generate novel motion combinations that have never been seen before. We can reach state-of-the-art render quality within just 5 minutes of training and in less than half an hour, we can synthesize novel views of dynamic scenes with superior photorealistic quality. Our representation is interpretable, efficient, and expressive enough to offer real-time view synthesis of complex dynamic scene motions, in monocular and multi-view scenarios.
翻译:摘要:由于时间动态性与运动复杂性,准确高效地建模动态场景与运动被视为一项极具挑战性的任务。为解决这些问题,我们提出DynMF——一种紧凑高效的表示方法,将动态场景分解为少量神经轨迹。我们认为,动态场景中各点的运动可分解为一小组显式或学习到的轨迹。我们精心设计的神经框架仅包含少量在时间维度上查询的基函数,可实现与3D高斯泼溅相当的渲染速度(超过120 FPS),同时存储量仅需静态场景的两倍。该神经表示对动态场景固有的欠约束运动场施加充分约束,从而实现高效快速的优化。具体而言,通过将每个点绑定至强制其共享基轨迹的运动系数,并谨慎地对运动系数施加稀疏性损失,我们能够解耦构成场景的独立运动、控制各运动成分,并生成前所未见的新颖运动组合。仅需5分钟训练即可达到当前最优的渲染质量;在不到半小时内,即可合成具有卓越逼真度的动态场景新视角。该表示兼具可解释性、高效性与表达力,足以在单目与多视角场景中实现复杂动态场景运动的实时视角合成。