Brain signal visualization has emerged as an active research area, serving as a critical interface between the human visual system and computer vision models. Although diffusion models have shown promise in analyzing functional magnetic resonance imaging (fMRI) data, including reconstructing high-quality images consistent with original visual stimuli, their accuracy in extracting semantic and silhouette information from brain signals remains limited. In this regard, we propose a novel approach, referred to as Controllable Mind Visual Diffusion Model (CMVDM). CMVDM extracts semantic and silhouette information from fMRI data using attribute alignment and assistant networks. Additionally, a residual block is incorporated to capture information beyond semantic and silhouette features. We then leverage a control model to fully exploit the extracted information for image synthesis, resulting in generated images that closely resemble the visual stimuli in terms of semantics and silhouette. Through extensive experimentation, we demonstrate that CMVDM outperforms existing state-of-the-art methods both qualitatively and quantitatively.
翻译:脑信号可视化已成为一个活跃的研究领域,它是人类视觉系统与计算机视觉模型之间的关键接口。尽管扩散模型在分析功能性磁共振成像(fMRI)数据方面展现出潜力,包括重建与原始视觉刺激一致的高质量图像,但其从脑信号中提取语义和轮廓信息的准确性仍有限。为此,我们提出一种名为可控心理视觉扩散模型(CMVDM)的新方法。CMVDM利用属性对齐和辅助网络从fMRI数据中提取语义和轮廓信息,并引入残差块以捕获超出语义和轮廓特征的信息。随后,我们借助控制模型充分利用所提取的信息进行图像合成,生成在语义和轮廓上与视觉刺激高度相似的图像。通过大量实验,我们证明CMVDM在定性和定量上均优于现有最先进方法。