Active measurements can be used to collect server characteristics on a large scale. This kind of metadata can help discovering hidden relations and commonalities among server deployments offering new possibilities to cluster and classify them. As an example, identifying a previously-unknown cybercriminal infrastructures can be a valuable source for cyber-threat intelligence. We propose herein an active measurement-based methodology for acquiring Transport Layer Security (TLS) metadata from servers and leverage it for their fingerprinting. Our fingerprints capture the characteristic behavior of the TLS stack primarily caused by the implementation, configuration, and hardware support of the underlying server. Using an empirical optimization strategy that maximizes information gain from every handshake to minimize measurement costs, we generated 10 general-purpose Client Hellos used as scanning probes to create a large database of TLS configurations used for classifying servers. We fingerprinted 28 million servers from the Alexa and Majestic toplists and two Command and Control (C2) blocklists over a period of 30 weeks with weekly snapshots as foundation for two long-term case studies: classification of Content Delivery Network and C2 servers. The proposed methodology shows a precision of more than 99 % and enables a stable identification of new servers over time. This study describes a new opportunity for active measurements to provide valuable insights into the Internet that can be used in security-relevant use cases.
翻译:主动测量可被用于大规模收集服务器特征。此类元数据有助于发现服务器部署间的隐藏关联与共性,为服务器聚类与分类提供新可能。例如,识别先前未知的网络犯罪基础设施可成为网络威胁情报的重要来源。本文提出一种基于主动测量的方法,用于从服务器获取传输层安全(TLS)元数据,并利用其进行指纹识别。我们的指纹特征主要捕捉由底层服务器的实现、配置及硬件支持导致的TLS协议栈特性行为。通过采用经验优化策略,最大化每次握手的信息增益以降低测量成本,我们生成了10个通用型Client Hello扫描探针,构建了用于服务器分类的大型TLS配置数据库。在为期30周的时间内,我们以周快照为基准,对Alexa与Majestic顶级域名列表及两个命令与控制(C2)黑名单中的2800万台服务器进行了指纹识别,并以此为基础开展两项长期案例研究:内容分发网络与C2服务器的分类。所提方法显示出超过99%的精确率,并能够随时间推移稳定识别新增服务器。本研究揭示了主动测量为互联网提供安全相关场景中关键洞察的新途径。