In the past five years, the use of generative and foundational AI systems has greatly improved the decoding of brain activity. Visual perception, in particular, can now be decoded from functional Magnetic Resonance Imaging (fMRI) with remarkable fidelity. This neuroimaging technique, however, suffers from a limited temporal resolution ($\approx$0.5 Hz) and thus fundamentally constrains its real-time usage. Here, we propose an alternative approach based on magnetoencephalography (MEG), a neuroimaging device capable of measuring brain activity with high temporal resolution ($\approx$5,000 Hz). For this, we develop an MEG decoding model trained with both contrastive and regression objectives and consisting of three modules: i) pretrained embeddings obtained from the image, ii) an MEG module trained end-to-end and iii) a pretrained image generator. Our results are threefold: Firstly, our MEG decoder shows a 7X improvement of image-retrieval over classic linear decoders. Second, late brain responses to images are best decoded with DINOv2, a recent foundational image model. Third, image retrievals and generations both suggest that high-level visual features can be decoded from MEG signals, although the same approach applied to 7T fMRI also recovers better low-level features. Overall, these results, while preliminary, provide an important step towards the decoding -- in real-time -- of the visual processes continuously unfolding within the human brain.
翻译:过去五年中,生成式与基础人工智能系统的应用极大地提升了脑活动解码的效能。其中,视觉感知已能从功能磁共振成像(fMRI)中以显著保真度进行解码。然而,该神经影像技术受限于有限的时间分辨率(≈0.5 Hz),从而根本上制约了其实时应用。本文提出一种基于脑磁图(MEG)的替代方案——该神经影像设备能以高时间分辨率(≈5,000 Hz)测量脑活动。为此,我们开发了一个融合对比学习与回归目标的MEG解码模型,其包含三个模块:i) 从图像中获取的预训练嵌入,ii) 端到端训练的MEG模块,iii) 预训练图像生成器。主要结果有三项:首先,相较于传统线性解码器,我们的MEG解码器在图像检索性能上提升了7倍;其次,DINOv2这一新兴基础图像模型对晚期脑响应解码效果最佳;最后,图像检索与生成结果均表明,MEG信号可解码高级视觉特征,但相同方法应用于7T fMRI时能更有效地恢复低级特征。总体而言,这些初步成果为实时解码人类大脑中持续进行的视觉过程迈出了重要一步。