A/B tests are the gold standard for evaluating digital experiences on the web. However, traditional "fixed-horizon" statistical methods are often incompatible with the needs of modern industry practitioners as they do not permit continuous monitoring of experiments. Frequent evaluation of fixed-horizon tests ("peeking") leads to inflated type-I error and can result in erroneous conclusions. We have released an experimentation service on the Adobe Experience Platform based on anytime-valid confidence sequences, allowing for continuous monitoring of the A/B test and data-dependent stopping. We demonstrate how we adapted and deployed asymptotic confidence sequences in a full featured A/B testing platform, describe how sample size calculations can be performed, and how alternate test statistics like "lift" can be analyzed. On both simulated data and thousands of real experiments, we show the desirable properties of using anytime-valid methods instead of traditional approaches.
翻译:A/B测试是评估网络数字体验的黄金标准。然而,传统的“固定时域”统计方法通常与现代产业实践者的需求不兼容,因为它们不允许对实验进行持续监控。对固定时域测试的频繁评估(“偷窥”)会导致I型错误膨胀,并可能得出错误结论。我们在Adobe Experience Platform上发布了一项基于随时有效置信序列的实验服务,支持对A/B测试进行持续监控以及数据依赖的停止。我们展示了如何在功能完备的A/B测试平台中调整和部署渐近置信序列,描述了如何执行样本量计算,以及如何分析“提升度”等替代检验统计量。通过模拟数据和数千个真实实验,我们证明了使用随时有效方法相较于传统方法具有理想特性。