Network foundation models promise reusable representations for diverse traffic analysis tasks, but recent diagnostic works have revealed fundamental problems: models exploit dataset shortcuts rather than learning genuine traffic patterns, produce collapsed embedding spaces, and fail to capture the exogenous network conditions that shape real-world behavior. We translate these diagnostic insights into four concrete design principles: protocol-aware tokenization, operational context embedding, burst-flow hierarchical attention, and privacy-by-construction input design, and build netFound, a network foundation model whose architecture is motivated by this failure analysis. We pretrain netFound on a billion-token-scale corpus over 5000 GPU hours, and demonstrate that it produces high-quality representations with lower anisotropy, significantly higher alignment with domain-expert features, and an F1 of 0.95 on exogenous context discrimination where existing state-of-the-art models score below 0.62, while preserving privacy by excluding payload and IP addresses. netFound demonstrates significant improvements in frozen-encoder evaluation, showing that pretrained embeddings themselves carry useful structure, and remains the top performer across all benchmarks in end-to-end fine-tuned settings. We release full open-source code, weights for three model sizes on HuggingFace, a containerized pipeline from raw PCAPs to downstream inference, and the full 4.2 billion flows pretraining dataset to facilitate reproducibility and further research.


翻译:网络基础模型有望为多样化的流量分析任务提供可复用的表示,但近期诊断性工作揭示了根本性问题:模型利用数据集捷径而非学习真实流量模式,产生坍塌的嵌入空间,并且无法捕捉塑造真实世界行为的外生网络条件。我们将这些诊断性见解转化为四项具体设计原则——协议感知分词化、运行环境上下文嵌入、突发流层次注意力,以及自带隐私保护的输入设计,并基于此故障分析构建了网络基础模型netFound。我们在超过5000 GPU小时、十亿级token的语料上预训练netFound,证明其能生成高质量表示:各向异性更低,与领域专家特征的对齐显著增强,在外生性上下文判别任务中F1值达0.95(现有最优模型低于0.62),同时通过排除载荷和IP地址实现隐私保护。netFound在冻结编码器评估中展现显著改进,表明预训练嵌入本身已携带有效结构,并在端到端微调设置下保持所有基准测试的最佳性能。我们发布完整开源代码、HuggingFace上三种模型规模的权重、从原始PCAP文件到下游推理的容器化流水线,以及完整的42亿条流预训练数据集,以促进可复现性和后续研究。

0
下载
关闭预览

相关内容

ACM/IEEE第23届模型驱动工程语言和系统国际会议,是模型驱动软件和系统工程的首要会议系列,由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来,模型涵盖了建模的各个方面,从语言和方法到工具和应用程序。模特的参加者来自不同的背景,包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛,参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会,并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。 官网链接:http://www.modelsconference.org/
异质信息网络分析与应用综述,软件学报-北京邮电大学
神经网络的拓扑结构,TOPOLOGY OF DEEP NEURAL NETWORKS
专知会员服务
35+阅读 · 2020年4月15日
用户画像基础
DataFunTalk
12+阅读 · 2020年8月1日
网络表示学习概述
机器学习与推荐算法
20+阅读 · 2020年3月27日
PointNet系列论文解读
人工智能前沿讲习班
17+阅读 · 2019年5月3日
网络表示学习介绍
人工智能前沿讲习班
18+阅读 · 2018年11月26日
国家自然科学基金
9+阅读 · 2017年12月31日
国家自然科学基金
2+阅读 · 2017年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
8+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
VIP会员
最新内容
俄乌无人机战争的六大启示
专知会员服务
8+阅读 · 8月3日
《无人机空中监控:通信实验洞察》
专知会员服务
6+阅读 · 8月3日
从采集到决策:美军视角下的战术情报范式重构
《履带式无人地面战车技术发展现状》
专知会员服务
6+阅读 · 8月2日
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
7+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
14+阅读 · 7月31日
相关VIP内容
异质信息网络分析与应用综述,软件学报-北京邮电大学
神经网络的拓扑结构,TOPOLOGY OF DEEP NEURAL NETWORKS
专知会员服务
35+阅读 · 2020年4月15日
相关资讯
用户画像基础
DataFunTalk
12+阅读 · 2020年8月1日
网络表示学习概述
机器学习与推荐算法
20+阅读 · 2020年3月27日
PointNet系列论文解读
人工智能前沿讲习班
17+阅读 · 2019年5月3日
网络表示学习介绍
人工智能前沿讲习班
18+阅读 · 2018年11月26日
相关基金
国家自然科学基金
9+阅读 · 2017年12月31日
国家自然科学基金
2+阅读 · 2017年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
8+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
Top
微信扫码咨询专知VIP会员