Time series forecasting drives operational decisions in areas like finance, transportation, and energy. While supervised learning approaches achieve strong performance, they require domain-specific training, feature engineering, and ongoing maintenance. Large-scale foundation models have recently emerged as a zero-shot alternative, avoiding task-specific training much like LLMs. In this work, we evaluate foundation models against standard supervised approaches. Rather than focusing solely on aggregate accuracy, we analyze performance across four operational regimes: periodic human-centric systems, physically constrained processes, stochastic financial markets, and heterogeneous demand forecasting. Our results characterize optimal deployment areas. Foundation models perform well in domains with transferable periodic structures and are efficient for cold-start or long-tail scenarios. Conversely, supervised specialists maintain higher precision in systems governed by strict physical constraints. In financial domains, newer foundation models are rapidly closing the performance gap with supervised specialists. We further quantify trade-offs in inference latency, data drift adaptability, and deployment constraints. Finally, we propose a Complexity Router that assigns each series to the optimal model class using empirical features. We demonstrate that this selective routing achieves higher accuracy and significantly lower inference costs compared to deploying a universal foundation model, providing a practical framework for balancing generalization and efficiency.
翻译:时间序列预测驱动着金融、交通和能源等领域的操作决策。尽管监督学习方法性能强劲,但它们需要特定领域的训练、特征工程及持续维护。近期,大规模基础模型作为零样本替代方案出现,如同大语言模型一般避开了任务特定训练。本研究将基础模型与标准监督方法进行对比评估。我们并非仅关注总体准确率,而是在四种操作情境下分析其性能:周期性人本系统、物理约束过程、随机性金融市场及异质性需求预测。结果揭示了最佳部署领域:基础模型在具有可迁移周期结构的领域中表现优异,且适用于冷启动或长尾场景;相反,监督专家模型在严格物理约束主导的系统中保持更高精度。在金融领域,新一代基础模型正快速缩小与监督专家模型的性能差距。我们进一步量化了推理延迟、数据漂移适应性及部署约束间的权衡。最后,提出一种复杂度路由器,该路由器利用经验特征将各序列分配给最优模型类别。实验证明,与部署通用基础模型相比,这种选择性路由实现了更高准确率与显著更低的推理成本,为平衡泛化性与效率提供了实用框架。