In recent years, \emph{learned cardinality estimation} has emerged as an alternative to traditional query optimization methods: by training machine learning models over observed query performance, learned cardinality estimation techniques can accurately predict query cardinalities and costs -- accounting for skew, correlated predicates, and many other factors that traditional methods struggle to capture. However, query-driven learned cardinality estimators are dependent on sample workloads, requiring vast amounts of labeled queries. Further, we show that state-of-the-art query-driven techniques can make significant and unpredictable errors on queries that are outside the distribution of their training set. We show that these out-of-distribution errors can be mitigated by incorporating the \emph{domain knowledge} used in traditional query optimizers: \emph{constraints} on values and cardinalities (e.g., based on key-foreign-key relationships, range predicates, and more generally on inclusion and functional dependencies). We develop methods for \emph{semi-supervised} query-driven learned query optimization, based on constraints, and we experimentally demonstrate that such techniques can increase a learned query optimizer's accuracy in cardinality estimation, reduce the reliance on massive labeled queries, and improve the robustness of query end-to-end performance.
翻译:近年来,学习型基数估计作为传统查询优化方法的替代方案应运而生:通过在观察到的查询性能上训练机器学习模型,学习型基数估计技术能够准确预测查询基数和代价——涵盖传统方法难以捕捉的偏斜、相关谓词及其他诸多因素。然而,查询驱动的学习型基数估计器依赖于样本工作负载,需要大量标注查询。进一步地,我们发现最先进的查询驱动技术可能对训练集分布范围外的查询产生显著且不可预测的误差。我们表明,通过引入传统查询优化器中使用的领域知识(即基于键外键关系、范围谓词以及更广义的包含依赖和函数依赖的值与基数约束),可以缓解这些分布外误差。我们开发了基于约束的半监督查询驱动学习型查询优化方法,并通过实验证明:此类技术能提升学习型查询优化器在基数估计中的准确性,减少对海量标注查询的依赖,同时增强查询端到端性能的鲁棒性。