Reconstructing perceived natural images or decoding their categories from fMRI signals are challenging tasks with great scientific significance. Due to the lack of paired samples, most existing methods fail to generate semantically recognizable reconstruction and are difficult to generalize to novel classes. In this work, we propose, for the first time, a task-agnostic brain decoding model by unifying the visual stimulus classification and reconstruction tasks in a semantic space. We denote it as BrainCLIP, which leverages CLIP's cross-modal generalization ability to bridge the modality gap between brain activities, images, and texts. Specifically, BrainCLIP is a VAE-based architecture that transforms fMRI patterns into the CLIP embedding space by combining visual and textual supervision. Note that previous works rarely use multi-modal supervision for visual stimulus decoding. Our experiments demonstrate that textual supervision can significantly boost the performance of decoding models compared to the condition where only image supervision exists. BrainCLIP can be applied to multiple scenarios like fMRI-to-image generation, fMRI-image-matching, and fMRI-text-matching. Compared with BraVL, a recently proposed multi-modal method for fMRI-based brain decoding, BrainCLIP achieves significantly better performance on the novel class classification task. BrainCLIP also establishes a new state-of-the-art for fMRI-based natural image reconstruction in terms of high-level image features.
翻译:从fMRI信号中重建感知到的自然图像或解码其类别是具有重大科学意义的挑战性任务。由于缺乏配对样本,现有多数方法难以生成语义可辨识的重建结果,且难以泛化至新类别。本文首次提出一种任务无关的大脑解码模型,在语义空间中统一了视觉刺激分类与重建任务。我们将其命名为BrainCLIP,该模型利用CLIP的跨模态泛化能力,弥合大脑活动、图像和文本之间的模态鸿沟。具体而言,BrainCLIP采用基于VAE的架构,通过融合视觉与文本监督将fMRI模式映射至CLIP嵌入空间。值得注意的是,以往研究极少在视觉刺激解码中使用多模态监督。实验表明,与仅使用图像监督的情况相比,文本监督能够显著提升解码模型的性能。BrainCLIP可应用于多种场景,如fMRI到图像生成、fMRI-图像匹配及fMRI-文本匹配。与近期提出的基于fMRI大脑解码的多模态方法BraVL相比,BrainCLIP在新类别分类任务上取得了显著更优的性能。同时,BrainCLIP在基于fMRI的自然图像重建任务中,于高层级图像特征方面创下了新的最优纪录。