Deciphering visual content from functional Magnetic Resonance Imaging (fMRI) helps illuminate the human vision system. However, the scarcity of fMRI data and noise hamper brain decoding model performance. Previous approaches primarily employ subject-specific models, sensitive to training sample size. In this paper, we explore a straightforward but overlooked solution to address data scarcity. We propose shallow subject-specific adapters to map cross-subject fMRI data into unified representations. Subsequently, a shared deeper decoding model decodes cross-subject features into the target feature space. During training, we leverage both visual and textual supervision for multi-modal brain decoding. Our model integrates a high-level perception decoding pipeline and a pixel-wise reconstruction pipeline guided by high-level perceptions, simulating bottom-up and top-down processes in neuroscience. Empirical experiments demonstrate robust neural representation learning across subjects for both pipelines. Moreover, merging high-level and low-level information improves both low-level and high-level reconstruction metrics. Additionally, we successfully transfer learned general knowledge to new subjects by training new adapters with limited training data. Compared to previous state-of-the-art methods, notably pre-training-based methods (Mind-Vis and fMRI-PTE), our approach achieves comparable or superior results across diverse tasks, showing promise as an alternative method for cross-subject fMRI data pre-training. Our code and pre-trained weights will be publicly released at https://github.com/YulongBonjour/See_Through_Their_Minds.
翻译:从功能性磁共振成像(fMRI)中解码视觉内容有助于揭示人类视觉系统。然而,fMRI数据的稀缺性和噪声阻碍了脑解码模型的性能。以往方法主要采用对被试个体特异性建模的模型,这类模型对训练样本量较为敏感。本文探讨了一种直观但被忽视的解决方案来应对数据稀缺问题:我们提出使用浅层被试特异性适配器,将跨被试fMRI数据映射为统一表征;随后,共享的深层解码模型将跨被试特征解码至目标特征空间。训练过程中,我们同时利用视觉和文本监督信号进行多模态脑解码。模型融合了高级感知解码管道与受高级感知引导的像素级重建管道,模拟了神经科学中的自下而上与自上而下处理过程。实验表明,两个管道在跨被试神经表征学习上均具有鲁棒性。此外,融合高低层信息能同步提升低层与高层重建指标。我们还成功将学得的通用知识迁移至新被试:通过使用有限训练数据训练新适配器即实现迁移。与现有最先进方法(特别是基于预训练的方法Mind-Vis和fMRI-PTE)相比,本方法在多种任务中取得可比或更优结果,展现出作为跨被试fMRI数据预训练替代方法的潜力。我们的代码与预训练权重将在https://github.com/YulongBonjour/See_Through_Their_Minds 公开提供。