Recent progress in brain-guided image generation has improved the quality of fMRI-based reconstructions; however, fundamental challenges remain in preserving object-level structure and semantic fidelity. Many existing approaches overlook the spatial arrangement of salient objects, leading to conceptually inconsistent outputs. We propose a saliency-driven decoding framework that employs graph-informed saliency priors to translate structural cues from brain signals into spatial masks. These masks, together with semantic information extracted from embeddings, condition a diffusion model to guide image regeneration, helping preserve object conformity while maintaining natural scene composition. In contrast to pipelines that invoke multiple diffusion stages, our approach relies on a single frozen model, offering a more lightweight yet effective design. Experiments show that this strategy improves both conceptual alignment and structural similarity to the original stimuli, while also introducing a new direction for efficient, interpretable, and structurally grounded brain decoding.
翻译:近期,脑引导图像生成的进展提升了基于功能磁共振成像(fMRI)重建的质量,但在保留对象级结构和语义保真度方面仍存在根本性挑战。现有方法多忽略显著对象的空间排布,导致生成结果在概念层面存在不一致性。我们提出一种显著性引导的解码框架,通过引入基于图的显著性先验,将脑信号中的结构线索转化为空间掩码。这些掩码与嵌入中提取的语义信息共同调控扩散模型,引导图像再生过程,在保持自然场景构图的同时维护对象一致性。相较于需要多阶段扩散处理的流水线方法,本方法仅依赖单一冻结模型,以更轻量化的设计实现高效解码。实验表明,该策略在概念对齐度与原始刺激的结构相似性方面均取得提升,并为构建高效、可解释且具有结构根基的脑解码方案开辟了新方向。