Humans have a strong intuitive understanding of physical processes such as fluid falling by just a glimpse of such a scene picture, i.e., quickly derived from our immersive visual experiences in memory. This work achieves such a photo-to-fluid-dynamics reconstruction functionality learned from unannotated videos, without any supervision of ground-truth fluid dynamics. In a nutshell, a differentiable Euler simulator modeled with a ConvNet-based pressure projection solver, is integrated with a volumetric renderer, supporting end-to-end/coherent differentiable dynamic simulation and rendering. By endowing each sampled point with a fluid volume value, we derive a NeRF-like differentiable renderer dedicated from fluid data; and thanks to this volume-augmented representation, fluid dynamics could be inversely inferred from the error signal between the rendered result and ground-truth video frame (i.e., inverse rendering). Experiments on our generated Fluid Fall datasets and DPI Dam Break dataset are conducted to demonstrate both effectiveness and generalization ability of our method.
翻译:人类仅凭瞥见场景图片的一瞬间,便能对诸如流体下落等物理过程产生强烈的直觉理解,这源自记忆中沉浸式视觉经验的快速推导。本研究实现了从无标注视频中学习照片到流体动力学重建的功能,无需任何真实流体动力学标注。简而言之,我们构建了一个可微欧拉模拟器,其中采用基于卷积网络的压力投影求解器,并将其与体积渲染器集成,从而支持端到端/连贯的可微动态模拟与渲染。通过为每个采样点赋予流体体积值,我们推导出一种专为流体数据设计的类NeRF可微渲染器;借助这种体积增强表示,流体动力学能够从渲染结果与真实视频帧之间的误差信号(即逆渲染)中反向推断。我们在生成的Fluid Fall数据集和DPI Dam Break数据集上进行了实验,验证了该方法的效果和泛化能力。