We present TSNBench, the first benchmark for evaluating large language model (LLM) proficiency in Time-Sensitive Networking (TSN), a suite of IEEE 802.1 standards for deterministic communication with bounded latency in safety-critical domains such as autonomous vehicles, aviation, defense, and industrial automation. While LLMs have been extensively evaluated on general knowledge tasks, their capabilities in safety-critical networking domains remain largely unexplored. TSNBench comprises 939 expert-validated multiple-choice questions (MCQs) covering diverse TSN mechanisms, along with 100 open-ended Worst-Case Delay (WCD) computation tasks for Credit-Based Shaper (CBS) and Cyclic Queuing and Forwarding (CQF) across varying network topologies and traffic conditions. MCQ answers are validated by domain experts, and open-ended ground truth WCD values are computed using a verified Network Calculus (NC) solver for CBS and closed-form mathematical upper bounds for CQF. We evaluate 16 LLMs and find that although models achieve 67 to 95% accuracy on MCQs, they fail substantially on open-ended WCD computation. For CBS, only GPT-5 achieves a Mean Absolute Percentage Error (MAPE) of 36.2%, meaning its predicted WCD deviates by 36.2% of the actual TSN flow delay on average, while most models exceed 80%. For CQF, the best model achieves 41.8% MAPE, with most models clustering between 80% and 100%. Such errors are large relative to TSN latency budgets and can lead to violations of real-time constraints and unsafe configurations. TSNBench demonstrates that MCQ benchmarks may overestimate LLM capabilities in safety-critical networking domains.
翻译:我们提出TSNBench,这是首个用于评估大语言模型(LLM)在时间敏感网络(TSN)领域能力的基准测试。TSN是一套IEEE 802.1标准,用于在自动驾驶、航空、国防和工业自动化等安全关键领域实现具有有界延迟的确定性通信。尽管LLM已在通用知识任务上得到广泛评估,但其在安全关键网络领域的能力仍鲜有探索。TSNBench包含939道经专家验证的多选题(MCQ),涵盖多种TSN机制,以及100个开放式的基于信用整形器(CBS)和循环排队转发(CQF)的最坏情况延迟(WCD)计算任务,这些任务涉及不同网络拓扑和流量条件。多选题答案由领域专家验证,开放式的真实WCD值通过已验证的网络演算(NC)求解器(针对CBS)和封闭式数学上界(针对CQF)计算得出。我们评估了16个LLM,发现虽然模型在多选题上达到了67%到95%的准确率,但在开放式的WCD计算上表现显著不足。对于CBS,仅有GPT-5实现了36.2%的平均绝对百分比误差(MAPE),即其预测WCD平均偏离实际TSN流延迟的36.2%,而大多数模型误差超过80%。对于CQF,最佳模型实现了41.8%的MAPE,大多数模型误差集中在80%到100%之间。相对于TSN延迟预算,此类误差过大,可能导致违反实时约束并产生不安全配置。TSNBench表明,多选题基准测试可能高估LLM在安全关键网络领域的能力。