Speech brain-computer interfaces (BCIs), which translate brain signals into spoken words or sentences, have shown significant potential for high-performance BCI communication. Phonemes are the fundamental units of pronunciation in most languages. While existing speech BCIs have largely focused on English, where words contain diverse compositions of phonemes, Chinese Mandarin is a monosyllabic language, with words typically consisting of a consonant and a vowel. This feature makes it feasible to develop high-performance Mandarin speech BCIs by decoding phonemes directly from neural signals. This study aimed to decode spoken Mandarin phonemes using intracortical neural signals. We observed that phonemes with similar pronunciations were often represented by inseparable neural patterns, leading to confusion in phoneme decoding. This finding suggests that the neural representation of spoken phonemes has a hierarchical structure. To account for this, we proposed learning the neural representation of phoneme pronunciation in a hyperbolic space, where the hierarchical structure could be more naturally optimized. Experiments with intracortical neural signals from a Chinese participant showed that the proposed model learned discriminative and interpretable hierarchical phoneme representations from neural signals, significantly improving Chinese phoneme decoding performance and achieving state-of-the-art. The findings demonstrate the feasibility of constructing high-performance Chinese speech BCIs based on phoneme decoding.
翻译:言语脑机接口(BCI)通过将脑信号转化为语音词汇或句子,在高性能BCI通信领域展现出显著潜力。音素是大多数语言中发音的基本单元。现有言语BCI主要聚焦于英语——其词汇包含多样化的音素组合,而中文普通话属于单音节语言,词汇通常由辅音与元音构成。这一特性使得通过直接从神经信号解码音素来开发高性能普通话BCI成为可能。本研究旨在利用脑皮层神经信号解码口语普通话音素。我们观察到,发音相似的音素常呈现难以分离的神经模式,导致音素解码混乱。这一发现表明口语音素的神经表征具有层级结构。为此,我们提出在双曲空间中学习音素发音的神经表征——该空间能更自然地优化层级结构。基于一位中国被试的脑皮层神经信号实验表明,所提模型可从神经信号中学习到可区分且可解释的层级音素表征,显著提升了中文音素解码性能并达到当前最优水平。这些发现证实了基于音素解码构建高性能中文言语BCI的可行性。