In the era of Artificial Intelligence Generated Content (AIGC), conditional multimodal synthesis technologies (e.g., text-to-image, text-to-video, text-to-audio, etc) are gradually reshaping the natural content in the real world. The key to multimodal synthesis technology is to establish the mapping relationship between different modalities. Brain signals, serving as potential reflections of how the brain interprets external information, exhibit a distinctive One-to-Many correspondence with various external modalities. This correspondence makes brain signals emerge as a promising guiding condition for multimodal content synthesis. Brian-conditional multimodal synthesis refers to decoding brain signals back to perceptual experience, which is crucial for developing practical brain-computer interface systems and unraveling complex mechanisms underlying how the brain perceives and comprehends external stimuli. This survey comprehensively examines the emerging field of AIGC-based Brain-conditional Multimodal Synthesis, termed AIGC-Brain, to delineate the current landscape and future directions. To begin, related brain neuroimaging datasets, functional brain regions, and mainstream generative models are introduced as the foundation of AIGC-Brain decoding and analysis. Next, we provide a comprehensive taxonomy for AIGC-Brain decoding models and present task-specific representative work and detailed implementation strategies to facilitate comparison and in-depth analysis. Quality assessments are then introduced for both qualitative and quantitative evaluation. Finally, this survey explores insights gained, providing current challenges and outlining prospects of AIGC-Brain. Being the inaugural survey in this domain, this paper paves the way for the progress of AIGC-Brain research, offering a foundational overview to guide future work.
翻译:在人工智能生成内容(AIGC)时代,条件性多模态合成技术(如文生图、文生视频、文生音频等)正逐步重塑现实世界中的自然内容。多模态合成技术的核心在于建立不同模态之间的映射关系。脑信号作为大脑解读外部信息的潜在反映,与多种外部模态呈现出独特的“一对多”对应关系。这种对应关系使得脑信号成为多模态内容合成中富有潜力的引导条件。脑条件多模态合成指将脑信号解码为感知体验,这对于开发实用的脑机接口系统以及揭示大脑感知和理解外部刺激的复杂机制至关重要。本综述全面审视了基于AIGC的脑条件多模态合成这一新兴领域(称为AIGC-Brain),以描绘当前研究格局与未来方向。首先,作为AIGC-Brain解码与分析的基础,本文介绍了相关的脑神经影像数据集、功能脑区及主流生成模型。其次,我们为AIGC-Brain解码模型提供了全面的分类体系,并展示了面向特定任务的代表性工作与详细实现策略,以促进比较与深入分析。随后,介绍了用于定性与定量评估的质量评价方法。最后,本综述探讨了已取得的洞见,总结了当前面临的挑战,并勾勒了AIGC-Brain的发展前景。作为该领域的首篇综述,本文为AIGC-Brain研究的进展铺平了道路,提供了基础性概览以指导未来工作。