We present TSNBench, the first benchmark for evaluating large language model (LLM) proficiency in Time-Sensitive Networking (TSN), a suite of IEEE 802.1 standards for deterministic communication with bounded latency in safety-critical domains such as autonomous vehicles, aviation, defense, and industrial automation. While LLMs have been extensively evaluated on general knowledge tasks, their capabilities in safety-critical networking domains remain largely unexplored. TSNBench comprises 939 expert-validated multiple-choice questions (MCQs) covering diverse TSN mechanisms, along with 100 open-ended Worst-Case Delay (WCD) computation tasks for Credit-Based Shaper (CBS) and Cyclic Queuing and Forwarding (CQF) across varying network topologies and traffic conditions. MCQ answers are validated by domain experts, and open-ended ground truth WCD values are computed using a verified Network Calculus (NC) solver for CBS and closed-form mathematical upper bounds for CQF. We evaluate 16 LLMs and find that although models achieve 67 to 95% accuracy on MCQs, they fail substantially on open-ended WCD computation. For CBS, only GPT-5 achieves a Mean Absolute Percentage Error (MAPE) of 36.2%, meaning its predicted WCD deviates by 36.2% of the actual TSN flow delay on average, while most models exceed 80%. For CQF, the best model achieves 41.8% MAPE, with most models clustering between 80% and 100%. Such errors are large relative to TSN latency budgets and can lead to violations of real-time constraints and unsafe configurations. TSNBench demonstrates that MCQ benchmarks may overestimate LLM capabilities in safety-critical networking domains.


翻译:我们提出TSNBench,这是首个用于评估大语言模型(LLM)在时间敏感网络(TSN)领域能力的基准测试。TSN是一套IEEE 802.1标准,用于在自动驾驶、航空、国防和工业自动化等安全关键领域实现具有有界延迟的确定性通信。尽管LLM已在通用知识任务上得到广泛评估,但其在安全关键网络领域的能力仍鲜有探索。TSNBench包含939道经专家验证的多选题(MCQ),涵盖多种TSN机制,以及100个开放式的基于信用整形器(CBS)和循环排队转发(CQF)的最坏情况延迟(WCD)计算任务,这些任务涉及不同网络拓扑和流量条件。多选题答案由领域专家验证,开放式的真实WCD值通过已验证的网络演算(NC)求解器(针对CBS)和封闭式数学上界(针对CQF)计算得出。我们评估了16个LLM,发现虽然模型在多选题上达到了67%到95%的准确率,但在开放式的WCD计算上表现显著不足。对于CBS,仅有GPT-5实现了36.2%的平均绝对百分比误差(MAPE),即其预测WCD平均偏离实际TSN流延迟的36.2%,而大多数模型误差超过80%。对于CQF,最佳模型实现了41.8%的MAPE,大多数模型误差集中在80%到100%之间。相对于TSN延迟预算,此类误差过大,可能导致违反实时约束并产生不安全配置。TSNBench表明,多选题基准测试可能高估LLM在安全关键网络领域的能力。

0
下载
关闭预览

相关内容

Networking:IFIP International Conferences on Networking。 Explanation:国际网络会议。 Publisher:IFIP。 SIT: http://dblp.uni-trier.de/db/conf/networking/index.html
基于LSTM深层神经网络的时间序列预测
论智
22+阅读 · 2018年9月4日
NetworkMiner - 网络取证分析工具
黑白之道
16+阅读 · 2018年6月29日
推荐|上交大推出Texygen:文本生成模型的基准测试平台
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
VIP会员
最新内容
反制无人机:乌克兰提供的五点启示
专知会员服务
9+阅读 · 9月23日
《各指挥层级均亟需红队能力》报告
专知会员服务
8+阅读 · 9月23日
《航电任务系统框架(FAMOS)》50页报告
专知会员服务
5+阅读 · 9月22日
《对抗行动中的人工智能与自主性》智库报告
专知会员服务
10+阅读 · 9月22日
《从数据到胜利:战争中的分析优势之争》
专知会员服务
12+阅读 · 9月22日
战争不仅需要机器人:人类仍不可或缺
专知会员服务
6+阅读 · 9月21日
《描绘美国防部创新基础设施的未来蓝图》100页
专知会员服务
12+阅读 · 9月21日
相关VIP内容
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员