The reliability of controlled experiments, or "A/B tests," can often be compromised due to the phenomenon of network interference, wherein the outcome for one unit is influenced by other units. To tackle this challenge, we propose a machine learning-based method to identify and characterize heterogeneous network interference. Our approach accounts for latent complex network structures and automates the task of "exposure mapping'' determination, which addresses the two major limitations in the existing literature. We introduce "causal network motifs'' and employ transparent machine learning models to establish the most suitable exposure mapping that reflects underlying network interference patterns. Our method's efficacy has been validated through simulations on two synthetic experiments and a real-world, large-scale test involving 1-2 million Instagram users, outperforming conventional methods such as design-based cluster randomization and analysis-based neighborhood exposure mapping. Overall, our approach not only offers a comprehensive, automated solution for managing network interference and improving the precision of A/B testing results, but it also sheds light on users' mutual influence and aids in the refinement of marketing strategies.
翻译:控制实验(即"A/B测试")的可靠性常因网络干扰现象而受损——即某个实验单元的结果会受到其他单元的影响。为应对这一挑战,我们提出一种基于机器学习的方法,用于识别并表征异质性网络干扰。该方法能够处理潜在复杂网络结构,并自动化"暴露映射"确定任务,从而解决现有文献中的两大局限。我们引入"因果网络基序"概念,并借助可解释的机器学习模型,构建最能反映潜在网络干扰模式的暴露映射。通过两项合成实验及一项涉及100-200万Instagram用户的大规模真实测试,我们验证了该方法的效果:其性能优于基于设计的聚类随机化与基于分析的邻域暴露映射等传统方法。总体而言,本方法不仅为管理网络干扰、提升A/B测试结果精准度提供了一套全面自动化解决方案,还能揭示用户间的相互影响,助力营销策略优化。