This work proposes a data driven learning model for the synthesis of keystroke biometric data. The proposed method is compared with two statistical approaches based on Universal and User-dependent models. These approaches are validated on the bot detection task, using the keystroke synthetic data to improve the training process of keystroke-based bot detection systems. Our experimental framework considers a dataset with 136 million keystroke events from 168 thousand subjects. We have analyzed the performance of the three synthesis approaches through qualitative and quantitative experiments. Different bot detectors are considered based on several supervised classifiers (Support Vector Machine, Random Forest, Gaussian Naive Bayes and a Long Short-Term Memory network) and a learning framework including human and synthetic samples. The experiments demonstrate the realism of the synthetic samples. The classification results suggest that in scenarios with large labeled data, these synthetic samples can be detected with high accuracy. However, in few-shot learning scenarios it represents an important challenge. Furthermore, these results show the great potential of the presented models.
翻译:本文提出了一种数据驱动的学习模型,用于合成击键生物特征数据。我们将所提方法与两种基于通用模型和用户依赖模型的统计方法进行了比较。这些方法在机器人检测任务中得到了验证,利用击键合成数据改进了基于击键行为的机器人检测系统的训练过程。我们的实验框架包含来自16.8万受试者的1.36亿次击键事件数据。通过定性和定量实验,我们分析了三种合成方法的性能。基于多种监督分类器(支持向量机、随机森林、高斯朴素贝叶斯和长短期记忆网络)以及包含人类与合成样本的学习框架,我们考虑了不同的机器人检测器。实验证明了合成样本的真实性。分类结果表明,在拥有大量标注数据的场景中,这些合成样本能够被高精度检测。然而,在少样本学习场景中,这构成了重大挑战。此外,这些结果也展现了所提模型的巨大潜力。