The source code of Function as a Service (FaaS) applications is constantly being refined. To detect if a source code change introduces a significant performance regression, the traditional benchmarking approach evaluates both the old and new function version separately using numerous artificial requests. In this paper, we describe a wrapper approach that enables the Randomized Multiple Interleaved Trials (RMIT) benchmark execution methodology in FaaS environments and use bootstrapping percentile intervals to derive more accurate confidence intervals of detected performance changes. We evaluate our approach using two public FaaS providers, an artificial performance issue, and several benchmark configuration parameters. We conclude that RMIT can shrink the width of confidence intervals in the results from 10.65% using the traditional approach to 0.37% using RMIT and thus enables a more fine-grained performance change detection.
翻译:函数即服务(FaaS)应用程序的源代码不断被优化。为检测源代码变更是否引入显著的性能退化,传统基准测试方法通过使用大量人工请求分别评估新旧函数版本。本文描述了一种包装方法,使得随机化多重交错试验(RMIT)基准测试执行方法能够在FaaS环境中实现,并利用自助法百分位置信区间为检测到的性能变化推导出更精确的置信区间。我们使用两个公共FaaS提供商、一个人为性能问题以及多个基准测试配置参数对提出的方法进行了评估。结论表明,RMIT可将结果中置信区间的宽度从传统方法的10.65%缩减至0.37%,从而实现更细粒度的性能变化检测。