In many real-world settings, image observations of freely rotating 3D rigid bodies may be available when low-dimensional measurements are not. However, the high-dimensionality of image data precludes the use of classical estimation techniques to learn the dynamics. The usefulness of standard deep learning methods is also limited, because an image of a rigid body reveals nothing about the distribution of mass inside the body, which, together with initial angular velocity, is what determines how the body will rotate. We present a physics-based neural network model to estimate and predict 3D rotational dynamics from image sequences. We achieve this using a multi-stage prediction pipeline that maps individual images to a latent representation homeomorphic to $\mathbf{SO}(3)$, computes angular velocities from latent pairs, and predicts future latent states using the Hamiltonian equations of motion. We demonstrate the efficacy of our approach on new rotating rigid-body datasets of sequences of synthetic images of rotating objects, including cubes, prisms and satellites, with unknown uniform and non-uniform mass distributions. Our model outperforms competing baselines on our datasets, producing better qualitative predictions and reducing the error observed for the state-of-the-art Hamiltonian Generative Network by a factor of 2.
翻译:在许多实际场景中,当无法获得低维测量时,自由旋转的三维刚体的图像观测数据往往可获取。然而,图像数据的高维性使得传统估计技术无法用于学习其动力学特性。标准深度学习方法的实用性同样受限,因为刚体图像无法揭示其内部质量分布——而这一分布与初始角速度共同决定了物体的旋转方式。我们提出了一种基于物理的神经网络模型,用于从图像序列中估计和预测三维旋转动力学。该方法采用多阶段预测流水线:将单张图像映射到与$\mathbf{SO}(3)$同胚的隐空间表征,从隐空间配对数据计算角速度,并利用哈密顿运动方程预测未来隐状态。我们在包含立方体、棱柱和卫星等旋转物体的合成图像序列新型旋转刚体数据集上验证了该方法的效果,这些物体具有未知的均匀与非均匀质量分布。我们的模型在数据集上优于竞争基线方法,不仅产生更优的定性预测结果,还将最先进的哈密顿生成网络误差降低了2倍。