Visual motion processing is essential for humans to perceive and interact with dynamic environments. Despite extensive research in cognitive neuroscience, image-computable models that can extract informative motion flow from natural scenes in a manner consistent with human visual processing have yet to be established. Meanwhile, recent advancements in computer vision (CV), propelled by deep learning, have led to significant progress in optical flow estimation, a task closely related to motion perception. Here we propose an image-computable model of human motion perception by bridging the gap between biological and CV models. Specifically, we introduce a novel two-stages approach that combines trainable motion energy sensing with a recurrent self-attention network for adaptive motion integration and segregation. This model architecture aims to capture the computations in V1-MT, the core structure for motion perception in the biological visual system, while providing the ability to derive informative motion flow for a wide range of stimuli, including complex natural scenes. In silico neurophysiology reveals that our model's unit responses are similar to mammalian neural recordings regarding motion pooling and speed tuning. The proposed model can also replicate human responses to a range of stimuli examined in past psychophysical studies. The experimental results on the Sintel benchmark demonstrate that our model predicts human responses better than the ground truth, whereas the state-of-the-art CV models show the opposite. Our study provides a computational architecture consistent with human visual motion processing, although the physiological correspondence may not be exact.
翻译:视觉运动处理对于人类感知和交互动态环境至关重要。尽管认知神经科学领域已有大量研究,但目前尚未建立能够以与人类视觉处理一致的方式从自然场景中提取信息性运动流的可计算图像模型。与此同时,近年来深度学习驱动的计算机视觉(CV)领域取得了显著进展,特别是在与运动感知密切相关的光流估计任务上。本文通过弥合生物模型与CV模型之间的差距,提出了一种人类运动感知的可计算图像模型。具体而言,我们引入了一种新颖的两阶段方法,该方法将可训练的运动能量感知与用于自适应运动整合和分离的循环自注意力网络相结合。该模型架构旨在捕捉生物视觉系统中运动感知核心结构V1-MT的计算特性,同时具备为包括复杂自然场景在内的各种刺激推导信息性运动流的能力。计算机神经生理学分析表明,我们的模型单元响应在运动池化和速度调谐方面与哺乳动物神经记录相似。所提出的模型还能复现过去心理物理学研究中人类对多种刺激的响应。Sintel基准测试的实验结果表明,我们的模型预测人类响应的准确性优于真实值,而最先进的CV模型则呈现相反结果。本研究提供了一种与人类视觉运动处理一致的计算架构,尽管其生理对应关系可能并非完全精确。