Distributed dataflow systems such as Apache Spark or Apache Flink enable parallel, in-memory data processing on large clusters of commodity hardware. Consequently, the appropriate amount of memory to allocate to the cluster is a crucial consideration. In this paper, we analyze the challenge of efficient resource allocation for distributed data processing, focusing on memory. We emphasize that in-memory processing with in-memory data processing frameworks can undermine resource efficiency. Based on the findings of our trace data analysis, we compile requirements towards an automated solution for efficient cluster resource allocation.
翻译:分布式数据流系统(如Apache Spark或Apache Flink)能够在由商用硬件构成的大规模集群上实现并行的内存数据处理。因此,为集群分配适当的内存容量成为一项关键考量。本文分析了面向分布式数据处理的高效资源分配挑战,重点关注内存。我们强调,使用内存数据处理框架进行的内存处理可能破坏资源效率。基于对跟踪数据分析的结果,我们归纳出实现高效集群资源分配的自动化解决方案所需满足的需求。