Decoding non-invasive brain recordings is pivotal for advancing our understanding of human cognition but faces challenges due to individual differences and complex neural signal representations. Traditional methods often require customized models and extensive trials, lacking interpretability in visual reconstruction tasks. Our framework integrates 3D brain structures with visual semantics using a Vision Transformer 3D. This unified feature extractor efficiently aligns fMRI features with multiple levels of visual embeddings, eliminating the need for subject-specific models and allowing extraction from single-trial data. The extractor consolidates multi-level visual features into one network, simplifying integration with Large Language Models (LLMs). Additionally, we have enhanced the fMRI dataset with diverse fMRI-image-related textual data to support multimodal large model development. Integrating with LLMs enhances decoding capabilities, enabling tasks such as brain captioning, complex reasoning, concept localization, and visual reconstruction. Our approach demonstrates superior performance across these tasks, precisely identifying language-based concepts within brain signals, enhancing interpretability, and providing deeper insights into neural processes. These advances significantly broaden the applicability of non-invasive brain decoding in neuroscience and human-computer interaction, setting the stage for advanced brain-computer interfaces and cognitive models.
翻译:解码非侵入性脑记录对于推进人类认知理解至关重要,但由于个体差异和复杂神经信号表征而面临挑战。传统方法通常需要定制化模型和大量试次,在视觉重建任务中缺乏可解释性。我们的框架利用Vision Transformer 3D将三维脑结构与视觉语义相整合。该统一特征提取器能高效对齐功能磁共振成像特征与多层级视觉嵌入,无需受试者特定模型,并支持从单试次数据中提取特征。该提取器将多层级视觉特征整合至单一网络,简化了与大型语言模型的集成。此外,我们通过多样化的功能磁共振成像-图像相关文本数据增强了功能磁共振成像数据集,以支持多模态大模型开发。与大型语言模型的集成增强了解码能力,实现了脑信号描述、复杂推理、概念定位和视觉重建等任务。我们的方法在这些任务中展现出卓越性能,能精准识别脑信号中基于语言的概念,提升可解释性,并为神经过程提供更深入的洞见。这些进展显著拓宽了非侵入性脑解码在神经科学和人机交互领域的适用性,为先进脑机接口和认知模型的发展奠定了基础。