Skyline queries typically search a Pareto-optimal set from a given data set to solve the corresponding multiobjective optimization problem. As the number of criteria increases, the skyline presumes excessive data items, which yield a meaningless result. To address this curse of dimensionality, we proposed a k-dominant skyline in which the number of skyline members was reduced by relaxing the restriction on the number of dimensions, considering the uncertainty of data. Specifically, each data item was associated with a probability of appearance, which represented the probability of becoming a member of the k-dominant skyline. As data items appear continuously in data streams, the corresponding k-dominant skyline may vary with time. Therefore, an effective and rapid mechanism of updating the k-dominant skyline becomes crucial. Herein, we proposed two time-efficient schemes, Middle Indexing (MI) and All Indexing (AI), for k-dominant skyline in distributed edge-computing environments, where irrelevant data items can be effectively excluded from the compute to reduce the processing duration. Furthermore, the proposed schemes were validated with extensive experimental simulations. The experimental results demonstrated that the proposed MI and AI schemes reduced the computation time by approximately 13% and 56%, respectively, compared with the existing method.
翻译:天际线查询通常从给定数据集中搜索帕累托最优集,以求解相应的多目标优化问题。随着准则数量的增加,天际线会预设过多数据项,导致结果失去意义。为解决这一维数灾难问题,我们提出了一种k-主导天际线,通过放宽维度限制并考虑数据的不确定性,减少了天际线成员数量。具体而言,每个数据项都与一个出现概率相关联,该概率表示其成为k-主导天际线成员的可能性。由于数据项在数据流中持续出现,相应的k-主导天际线可能随时间变化。因此,高效且快速的k-主导天际线更新机制变得至关重要。本文针对分布式边缘计算环境提出了两种时间高效方案——中间索引(MI)与全索引(AI),通过有效排除无关数据项缩短了处理时间。此外,通过大量实验仿真验证了所提方案的有效性。实验结果表明,与现有方法相比,所提MI和AI方案的计算时间分别减少了约13%和56%。