To handle the scarcity and heterogeneity of electroencephalography (EEG) data in Brain-Computer Interface (BCI) tasks, and to harness the vast public data, we propose Neuro-GPT, a foundation model consisting of an EEG encoder and a GPT model. The foundation model is pre-trained on a large-scale public EEG dataset, using a self-supervised task which learns how to reconstruct the masked chunk in EEG. We then fine-tune the foundation model on a Motor Imagery Classification task where only 9 subjects are available. Experiments demonstrated that applying foundation model can significantly improve classification performance compared to the model trained from scratch, which provides evidence for the advanced generalizability of foundation model and the ability to address the challenges of data scarcity and heterogeneity.
翻译:摘要:为解决脑机接口(BCI)任务中脑电图(EEG)数据的稀缺性和异质性,并充分利用海量公共数据,我们提出Neuro-GPT——一种由脑电图编码器和GPT模型组成的基础模型。该基础模型通过自监督学习任务在大型公开脑电图数据集上进行预训练,该任务旨在学习如何重建脑电信号中屏蔽的片段。随后,我们在仅含9名受试者的运动想象分类任务上对基础模型进行微调。实验证明,与从零开始训练的模型相比,应用基础模型可显著提升分类性能,这为基础模型的优异泛化能力及其解决数据稀缺性和异质性挑战的能力提供了实证支持。