Logs are an essential source of information for people to understand the running status of a software system. Due to the evolving modern software architecture and maintenance methods, more research efforts have been devoted to automated log analysis. In particular, machine learning (ML) has been widely used in log analysis tasks. In ML-based log analysis tasks, converting textual log data into numerical feature vectors is a critical and indispensable step. However, the impact of using different log representation techniques on the performance of the downstream models is not clear, which limits researchers and practitioners' opportunities of choosing the optimal log representation techniques in their automated log analysis workflows. Therefore, this work investigates and compares the commonly adopted log representation techniques from previous log analysis research. Particularly, we select six log representation techniques and evaluate them with seven ML models and four public log datasets (i.e., HDFS, BGL, Spirit and Thunderbird) in the context of log-based anomaly detection. We also examine the impacts of the log parsing process and the different feature aggregation approaches when they are employed with log representation techniques. From the experiments, we provide some heuristic guidelines for future researchers and developers to follow when designing an automated log analysis workflow. We believe our comprehensive comparison of log representation techniques can help researchers and practitioners better understand the characteristics of different log representation techniques and provide them with guidance for selecting the most suitable ones for their ML-based log analysis workflow.
翻译:日志是了解软件系统运行状态的重要信息源。随着现代软件架构与维护方法的不断演进,自动化日志分析研究日益受到关注。特别是机器学习已广泛应用于日志分析任务中。在基于机器学习的日志分析过程中,将文本形式的日志数据转化为数值特征向量是至关重要的步骤。然而,不同日志表示技术对下游模型性能的影响尚不明确,这限制了研究人员和实践者在自动化日志分析工作流中选择最优日志表示技术的可能性。因此,本研究系统梳理并对比了以往日志分析研究中常用的日志表示技术。具体而言,我们选取六种日志表示技术,在基于日志的异常检测场景下,结合七种机器学习模型与四个公开日志数据集(HDFS、BGL、Spirit和Thunderbird)进行评估。同时,我们还探究了日志解析过程及不同特征聚合方法对日志表示技术应用效果的影响。通过实验,我们为未来研究者和开发者在设计自动化日志分析工作流时提供了若干启发式指导原则。我们相信,对日志表示技术的全面比较能帮助研究人员和实践者深入理解不同日志表示技术的特性,并为其在基于机器学习的日志分析工作流中选择最合适的技术提供参考依据。