We introduce an on-ground Pedestrian World Model, a computational model that can predict how pedestrians move around an observer in the crowd on the ground plane, but from just the egocentric-views of the observer. Our model, InCrowdFormer, fully leverages the Transformer architecture by modeling pedestrian interaction and egocentric to top-down view transformation with attention, and autoregressively predicts on-ground positions of a variable number of people with an encoder-decoder architecture. We encode the uncertainties arising from unknown pedestrian heights with latent codes to predict the posterior distributions of pedestrian positions. We validate the effectiveness of InCrowdFormer on a novel prediction benchmark of real movements. The results show that InCrowdFormer accurately predicts the future coordination of pedestrians. To the best of our knowledge, InCrowdFormer is the first-of-its-kind pedestrian world model which we believe will benefit a wide range of egocentric-view applications including crowd navigation, tracking, and synthesis.
翻译:摘要:我们提出了一种地面行人世界模型,这是一种计算模型,能够仅从观察者的自我中心视角预测人群中的行人如何在地面上围绕观察者移动。我们的模型InCrowdFormer充分利用Transformer架构,通过注意力机制建模行人交互以及从自我中心视角到俯视图的变换,并采用编码器-解码器结构自回归地预测可变数量人群在地面上的位置。我们通过潜在编码对未知行人高度引起的不确定性进行建模,以预测行人位置的后验分布。我们在一个包含真实运动轨迹的新型预测基准上验证了InCrowdFormer的有效性。结果表明,InCrowdFormer能够准确预测行人的未来协调行为。据我们所知,InCrowdFormer是首个此类行人世界模型,我们相信它将有益于一系列自我中心视角应用,包括人群导航、跟踪和合成。