Network operators monitor their infrastructure by collecting telemetry data such as packet counts, byte rates, or flow volumes, yet answering the questions that effective operations demand -- forecasting future load, diagnosing and characterizing anomalies, and searching for and retrieving historical precedents -- requires more than raw measurements. Bridging this gap calls for learned representations: compact per-entity summaries that capture temporal dynamics from each entity's univariate time series. Time-series foundation models are the natural starting point, but they are designed for dense, periodic benchmark datasets -- the \emph{mild} statistical regime. However, network telemetry data inhabits the \emph{wild} regime: operationally relevant events are rare, separated by variable-length stretches of low or no activity (``ebbs''), with intermittent bursts of heavy-tailed extremes (``tides''). We present NetBurst, an event-centric pipeline that collapses ebbs, separates each time series into a stream of burst timings and a stream of burst magnitudes, and learns a single representation serving all three operational tasks. Compared to the strongest competitors among eight baselines -- including Amazon's Chronos-2 and Datadog's Toto -- and across nine production telemetry configurations, NetBurst reduces median forecasting error by $1.3$--$116\times$ on wild-regime data with a $1.0$--$7.5\times$ better match to the true burst distribution, and matches baselines on mild-regime benchmarks. For characterizing anomalies, NetBurst produces balanced, well-spread clusters that are $16\times$ more describable in operator-familiar terms under a novel interpretability score, and cluster-filtered search delivers $7.5\times$ faster end-to-end retrieval.
翻译:摘要:网络运营商通过收集遥测数据(如数据包计数、字节速率或流量体积)来监控其基础设施,但有效运维所需解决的问题——预测未来负载、诊断和表征异常、搜索和检索历史先例——仅靠原始测量值是不够的。弥补这一差距需要学习表征:紧凑的逐实体摘要,从每个实体的单变量时间序列中捕获时间动态。时间序列基础模型是自然的起点,但它们是为密集、周期性的基准数据集设计的——即“温和”统计区间。然而,网络遥测数据处于“野外”区间:操作相关的事件稀少,被低活动或无活动(“退潮”)的可变长度间隔分隔,并伴有重尾极值(“潮汐”)的间歇性突发。我们提出NetBurst,一个事件驱动的流水线,它压缩退潮,将每个时间序列分离成突发时序流和突发幅度流,并学习一个服务于所有三种操作任务的单一表征。与八种基线中包括亚马逊Chronos-2和Datadog Toto在内的最强竞争对手相比,在九种生产遥测配置上,NetBurst在野外区间数据上将中位预测误差降低了$1.3$-$116$倍,与真实突发分布的匹配度提高了$1.0$-$7.5$倍,并在温和区间基准测试中匹配基线。对于异常表征,NetBurst生成平衡且分布良好的聚类,在新颖的可解释性评分下,以运维人员熟悉的术语描述性提高了$16$倍,且聚类过滤搜索实现了$7.5$倍的端到端检索加速。