The analysis of logs is a vital activity undertaken for fault or cyber incident detection, investigation and technical forensics analysis for system and cyber resilience. The potential application of AI algorithms for Log analysis could augment such complex and laborious tasks. However, such solution has its constraints the heterogeneity of log sources and limited to no labels for training a classifier. When such labels become available, the need for the classifier to be updated. This practice-based research seeks to address these challenges with the use of Transformer construct to train a new model with only normal log entries. Log augmentation through multiple forms of perturbation is applied as a form of self-supervised training for feature learning. The model is further finetuned using a form of reinforcement learning with a limited set of label samples to mimic real-world situation with the availability of labels. The experimental results of our model construct show promise with comparative evaluation measurements paving the way for future practical applications.
翻译:日志分析是进行故障或网络事件检测、调查以及系统与网络韧性技术取证分析的关键活动。AI算法在日志分析中的潜在应用可能增强这些复杂且耗力的任务。然而,此类解决方案存在其局限性,包括日志源的异构性以及训练分类器所需的标签缺失或极为有限。即便标签可用时,分类器仍需持续更新。这项基于实践的研究旨在通过利用Transformer架构仅使用正常日志条目训练新模型来应对这些挑战。我们采用多种扰动形式的日志增强作为自监督特征学习训练方法,并通过强化学习机制结合少量标签样本对模型进行微调,以模拟标签可用时的真实场景。实验结果表明,该模型架构在对比评估指标上展现出潜力,为未来实际应用奠定了基础。