In the Machine Learning (ML) literature, a well-known problem is the Dataset Shift problem where, differently from the ML standard hypothesis, the data in the training and test sets can follow different probability distributions, leading ML systems toward poor generalisation performances. This problem is intensely felt in the Brain-Computer Interface (BCI) context, where bio-signals as Electroencephalographic (EEG) are often used. In fact, EEG signals are highly non-stationary both over time and between different subjects. To overcome this problem, several proposed solutions are based on recent transfer learning approaches such as Domain Adaption (DA). In several cases, however, the actual causes of the improvements remain ambiguous. This paper focuses on the impact of data normalisation, or standardisation strategies applied together with DA methods. In particular, using \textit{SEED}, \textit{DEAP}, and \textit{BCI Competition IV 2a} EEG datasets, we experimentally evaluated the impact of different normalization strategies applied with and without several well-known DA methods, comparing the obtained performances. It results that the choice of the normalisation strategy plays a key role on the classifier performances in DA scenarios, and interestingly, in several cases, the use of only an appropriate normalisation schema outperforms the DA technique.
翻译:在机器学习文献中,数据集漂移问题是一个公认难题:与机器学习标准假设不同,训练集和测试集中的数据可能遵循不同的概率分布,导致机器学习系统泛化性能低下。该问题在脑-机接口领域尤为突出——该领域常使用脑电图等生物信号。实际上,EEG 信号在时间跨度和个体间均呈现高度非平稳性。为解决该问题,现有多种解决方案基于领域自适应等迁移学习方法,但在许多情况下,性能提升的实际原因仍不明确。本文聚焦数据标准化策略或与 DA 方法联合使用的标准化方法所产生的影响。具体而言,我们采用 SEED、DEAP 和 BCI Competition IV 2a 三个 EEG 数据集,系统评估了不同标准化策略在有无多种经典 DA 方法配合下的效果差异。实验结果表明:标准化策略的选择对 DA 场景下的分类器性能具有关键影响;值得关注的是,在多种情况下,仅采用恰当的标准化方案即可取得优于 DA 技术的效果。