Website Fingerprinting (WF) is considered a major threat to the anonymity of Tor users (and other anonymity systems). While state-of-the-art WF techniques have claimed high attack accuracies, e.g., by leveraging Deep Neural Networks (DNN), several recent works have questioned the practicality of such WF attacks in the real world due to the assumptions made in the design and evaluation of these attacks. In this work, we argue that such impracticality issues are mainly due to the attacker's inability in collecting training data in comprehensive network conditions, e.g., a WF classifier may be trained only on samples collected on specific high-bandwidth network links but deployed on connections with different network conditions. We show that augmenting network traces can enhance the performance of WF classifiers in unobserved network conditions. Specifically, we introduce NetAugment, an augmentation technique tailored to the specifications of Tor traces. We instantiate NetAugment through semi-supervised and self-supervised learning techniques. Our extensive open-world and close-world experiments demonstrate that under practical evaluation settings, our WF attacks provide superior performances compared to the state-of-the-art; this is due to their use of augmented network traces for training, which allows them to learn the features of target traffic in unobserved settings. For instance, with a 5-shot learning in a closed-world scenario, our self-supervised WF attack (named NetCLR) reaches up to 80% accuracy when the traces for evaluation are collected in a setting unobserved by the WF adversary. This is compared to an accuracy of 64.4% achieved by the state-of-the-art Triplet Fingerprinting [35]. We believe that the promising results of our work can encourage the use of network trace augmentation in other types of network traffic analysis.
翻译:网站指纹识别(WF)被认为是对Tor用户(及其他匿名系统)匿名性的主要威胁。尽管最先进的WF技术已声称取得高攻击准确率(例如通过利用深度神经网络),但近期多项研究质疑此类WF攻击在现实世界中的实用性,原因在于这些攻击的设计和评估中采用了诸多假设。本文指出,此类非实用性问题主要源于攻击者难以在全面网络条件下收集训练数据——例如,WF分类器可能仅基于特定高带宽网络链路采集的样本进行训练,却部署于不同网络条件的连接中。研究表明,增强网络流量可提升WF分类器在未观测网络条件下的性能。具体而言,我们提出NetAugment——一种针对Tor流量特性定制的增强技术,并通过半监督和自监督学习技术实现该框架。大量开放世界与封闭世界实验表明,在实际评估场景下,本方法因采用增强网络流量进行训练以学习未观测场景下目标流量特征,其攻击性能显著优于现有技术。例如,在封闭世界场景中采用5样本学习时,我们提出的自监督WF攻击方法NetCLR在评估流量采集于WF攻击者未观测环境下的情况下,准确率可达80%,而现有最优的Triplet Fingerprinting方法[35]仅能达到64.4%。我们相信,本研究的优秀成果将推动网络流量增强技术在其他网络流量分析领域的应用。