The popularity of Deep Learning (DL), coupled with network traffic visibility reduction due to the increased adoption of HTTPS, QUIC and DNS-SEC, re-ignited interest towards Traffic Classification (TC). However, to tame the dependency from task-specific large labeled datasets we need to find better ways to learn representations that are valid across tasks. In this work we investigate this problem comparing transfer learning, meta-learning and contrastive learning against reference Machine Learning (ML) tree-based and monolithic DL models (16 methods total). Using two publicly available datasets, namely MIRAGE19 (40 classes) and AppClassNet (500 classes), we show that (i) using large datasets we can obtain more general representations, (ii) contrastive learning is the best methodology and (iii) meta-learning the worst one, and (iv) while ML tree-based cannot handle large tasks but fits well small tasks, by means of reusing learned representations, DL methods are reaching tree-based models performance also for small tasks.
翻译:深度学习(DL)的普及,加之HTTPS、QUIC和DNS-SEC等协议的广泛采用导致网络流量可见性降低,重新激发了人们对流量分类(TC)的兴趣。然而,为摆脱对特定任务的大规模标注数据集的依赖,我们需要寻找更优方法来学习跨任务通用的表征。本研究通过将迁移学习、元学习和对比学习与基于决策树的传统机器学习(ML)模型及单体式DL模型(共16种方法)进行对比,探究该问题。基于两个公开数据集(MIRAGE19共40类,AppClassNet共500类),我们证明:(i)利用大规模数据集可获得更通用的表征;(ii)对比学习为最优方法;(iii)元学习方法表现最差;(iv)尽管基于决策树的ML方法无法处理大规模任务但适用于小规模任务,通过复用习得的表征,DL方法在小规模任务上的性能也已接近基于决策树的方法。