Decoding language from brain dynamics is an important open direction in the realm of brain-computer interface (BCI), especially considering the rapid growth of large language models. Compared to invasive-based signals which require electrode implantation surgery, non-invasive neural signals (e.g. EEG, MEG) have attracted increasing attention considering their safety and generality. However, the exploration is not adequate in three aspects: 1) previous methods mainly focus on EEG but none of the previous works address this problem on MEG with better signal quality; 2) prior works have predominantly used ``teacher-forcing" during generative decoding, which is impractical; 3) prior works are mostly ``BART-based" not fully auto-regressive, which performs better in other sequence tasks. In this paper, we explore the brain-to-text translation of MEG signals in a speech-decoding formation. Here we are the first to investigate a cross-attention-based ``whisper" model for generating text directly from MEG signals without teacher forcing. Our model achieves impressive BLEU-1 scores of 60.30 and 52.89 without pretraining \& teacher-forcing on two major datasets (\textit{GWilliams} and \textit{Schoffelen}). This paper conducts a comprehensive review to understand how speech decoding formation performs on the neural decoding tasks, including pretraining initialization, training \& evaluation set splitting, augmentation, and scaling law.
翻译:从大脑动态中解码语言是脑机接口领域的一个重要开放方向,尤其是在大型语言模型快速发展的背景下。相较于需要电极植入手术的侵入式信号,非侵入式神经信号(如脑电图、脑磁图)因其安全性和普适性而受到越来越多的关注。然而,当前研究在以下三个方面仍不充分:1)以往方法主要聚焦于脑电图,但尚无工作针对信号质量更优的脑磁图解决该问题;2)以往研究在生成解码过程中主要采用“教师强制”方法,这在实际应用中并不可行;3)以往研究大多基于“BART”架构,并非完全自回归,而后者在其他序列任务中表现更佳。本文探索了以语音解码方式从脑磁图信号进行脑到文本的翻译。我们首次研究了基于交叉注意力的“Whisper”模型,无需教师强制即可直接从脑磁图信号生成文本。在两大主要数据集(GWilliams和Schoffelen)上,我们的模型无需预训练及教师强制即取得了令人瞩目的BLEU-1分数:60.30和52.89。本文通过全面综述,深入理解了语音解码方式在神经解码任务中的表现,包括预训练初始化、训练与评估集划分、数据增强以及缩放规律。