Materialized views are a core construct in database systems, used to accelerate analytical queries and optimize batch pipelines for extract-transform-load (ETL) workflows. Maintaining view consistency as underlying data evolves is a fundamental challenge, especially in high-throughput and real-time settings. Incremental view maintenance (IVM) has been studied for decades and continues to attract significant investment from major database vendors. However, most industrial systems either offer limited SQL-operator coverage or require users to hand-tune refresh strategies. This paper presents Enzyme, an IVM engine developed at Databricks to power Spark Declarative Pipelines. It provides a built-in, end-to-end approach to incremental pipelines, utilizing materialized views as first-class building blocks. By automating refresh planning, Enzyme reduces total cost of ownership and lets users focus on business logic rather than MV mechanics. Validation across thousands of large-scale production pipelines spanning diverse application domains has demonstrated substantial computational efficiency gains, yielding a cumulative daily compute reduction of billions of CPU seconds. Built atop Apache Spark primitives, Enzyme adds a cost-based optimization layer that selects refresh strategies for collections of materialized views organized into pipelines. Enzyme's modular architecture is designed to generalize across data sources and query engines. We present key design decisions for incremental refresh planning and execution, including optimizations that exploit batching opportunities across materialized view sources. Experimental results on standard benchmarks demonstrate significant performance improvements at scale.


翻译:摘要:物化视图是数据库系统中的核心构造,用于加速分析查询并优化提取-转换-加载(ETL)工作流中的批处理管道。随着底层数据演变,维持视图一致性是一项基础性挑战,尤其是在高吞吐量和实时场景中。增量视图维护(IVM)已研究数十年,并持续吸引着主要数据库厂商的巨额投入。然而,大多数工业系统要么仅支持有限的SQL操作符覆盖范围,要么要求用户手动调整刷新策略。本文介绍了Enzyme——Databricks开发的IVM引擎,用于支持Spark声明式管道。它提供了面向增量管道的内置端到端方法,将物化视图作为一等构建块。通过自动化刷新规划,Enzyme降低了总拥有成本,使用户能够专注于业务逻辑而非物化视图的底层机制。在跨多个应用领域的数千条大规模生产管道上的验证表明,该系统取得了显著的计算效率提升,累计每日计算量减少了数十亿CPU秒。Enzyme基于Apache Spark原语构建,并添加了基于成本的优化层,为组织成管道的物化视图集合选择刷新策略。其模块化架构旨在跨数据源和查询引擎实现通用化。我们介绍了增量刷新规划与执行的关键设计决策,包括利用物化视图源间批处理机会的优化方案。在标准基准测试上的实验结果表明,该系统在大规模场景下实现了显著的性能提升。

0
下载
关闭预览

相关内容

在无标注条件下适配视觉—语言模型:全面综述
专知会员服务
13+阅读 · 2025年8月9日
图数据库综述
专知会员服务
18+阅读 · 2025年6月2日
以数据为中心的图机器学习
专知会员服务
38+阅读 · 2023年9月25日
【数据中台】什么是数据中台?
产业智能官
18+阅读 · 2019年7月30日
计算机视觉方向简介 | 视觉惯性里程计(VIO)
计算机视觉life
64+阅读 · 2019年6月16日
面试题:请简要介绍下tensorflow的计算图
七月在线实验室
14+阅读 · 2019年6月10日
领域应用 | 到底什么时候使用图数据库?
开放知识图谱
16+阅读 · 2019年4月19日
深度学习的目标检测技术演进:R-CNN、Fast R-CNN、Faster R-CNN
数据挖掘入门与实战
13+阅读 · 2018年4月6日
视觉里程计:起源、优势、对比、应用
计算机视觉life
18+阅读 · 2017年7月17日
国家自然科学基金
9+阅读 · 2017年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
Arxiv
15+阅读 · 2023年10月21日
VIP会员
最新内容
《无人机对海面作战影响评估》
专知会员服务
10+阅读 · 7月21日
印度精确打击与指挥架构的断层
专知会员服务
5+阅读 · 7月20日
美空军AI完成F-16战斗机自主空战历史性试飞
专知会员服务
6+阅读 · 7月20日
深入Project Maven:为何人工智能在战场上依然失灵
锻造未来士兵:外骨骼、基因工程与赛博格
专知会员服务
8+阅读 · 7月19日
相关基金
国家自然科学基金
9+阅读 · 2017年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员