Materialized model query aims to find the most appropriate materialized model as the initial model for model reuse. It is the precondition of model reuse, and has recently attracted much attention. {Nonetheless, the existing methods suffer from the need to provide source data, limited range of applications, and inefficiency since they do not construct a suitable metric to measure the target-related knowledge of materialized models. To address this, we present \textsf{MMQ}, a source-data free, general, efficient, and effective materialized model query framework.} It uses a Gaussian mixture-based metric called separation degree to rank materialized models. For each materialized model, \textsf{MMQ} first vectorizes the samples in the target dataset into probability vectors by directly applying this model, then utilizes Gaussian distribution to fit for each class of probability vectors, and finally uses separation degree on the Gaussian distributions to measure the target-related knowledge of the materialized model. Moreover, we propose an improved \textsf{MMQ} (\textsf{I-MMQ}), which significantly reduces the query time while retaining the query performance of \textsf{MMQ}. Extensive experiments on a range of practical model reuse workloads demonstrate the effectiveness and efficiency of \textsf{MMQ}.
翻译:物化模型查询旨在寻找最合适的物化模型,作为模型复用的初始模型。它是模型复用的前提条件,并近期受到广泛关注。然而,现有方法因未构建合适的度量标准来衡量物化模型的目标相关知识,存在需提供源数据、适用范围有限且效率低下的问题。为解决此问题,我们提出\textsf{MMQ}——一种无需源数据、通用、高效且有效的物化模型查询框架。该框架采用基于高斯混合的度量标准——分离度(separation degree)来对物化模型进行排序。对于每个物化模型,\textsf{MMQ}首先通过直接应用该模型,将目标数据集中的样本向量化为概率向量;随后利用高斯分布拟合每个类别的概率向量;最后基于高斯分布间的分离度度量物化模型的目标相关知识。此外,我们提出改进版\textsf{MMQ}(\textsf{I-MMQ}),在保持\textsf{MMQ}查询性能的同时显著降低查询时间。针对一系列实际模型复用工作负载的广泛实验,验证了\textsf{MMQ}的有效性与高效性。