As deep learning has achieved state-of-the-art performance for many tasks of EEG-based BCI, many efforts have been made in recent years trying to understand what have been learned by the models. This is commonly done by generating a heatmap indicating to which extent each pixel of the input contributes to the final classification for a trained model. Despite the wide use, it is not yet understood to which extent the obtained interpretation results can be trusted and how accurate they can reflect the model decisions. In order to fill this research gap, we conduct a study to evaluate different deep interpretation techniques quantitatively on EEG datasets. The results reveal the importance of selecting a proper interpretation technique as the initial step. In addition, we also find that the quality of the interpretation results is inconsistent for individual samples despite when a method with an overall good performance is used. Many factors, including model structure and dataset types, could potentially affect the quality of the interpretation results. Based on the observations, we propose a set of procedures that allow the interpretation results to be presented in an understandable and trusted way. We illustrate the usefulness of our method for EEG-based BCI with instances selected from different scenarios.
翻译:随着深度学习在基于脑电图(EEG)的脑机接口(BCI)多项任务中取得最优性能,近年来大量研究致力于理解这些模型所学习到的特征。通常的做法是通过生成热力图,展示输入图像中每个像素对训练模型最终分类结果的贡献程度。尽管该方法已被广泛采用,但目前尚不明确所获得的可解释结果的可靠程度及其反映模型决策的准确性。为填补这一研究空白,我们针对EEG数据集开展了定量评估不同深度可解释技术的研究。结果表明,选择恰当的可解释技术作为初始步骤至关重要。此外,我们发现即使使用整体性能良好的方法,单个样本的可解释结果质量仍存在不一致性。模型结构、数据集类型等多种因素均可能影响可解释结果的质量。基于这些观察,我们提出了一套流程,使可解释结果能够以可理解且可信赖的方式呈现。通过选取不同场景下的实例,我们验证了该方法在基于EEG的BCI系统中的实用性。