We cover how to determine a sufficiently large sample size for a $K$-armed randomized experiment in order to estimate conditional counterfactual expectations in data-driven subgroups. The sub-groups can be output by any feature space partitioning algorithm, including as defined by binning users having similar predictive scores or as defined by a learned policy tree. After carefully specifying the inference target, a minimum confidence level, and a maximum margin of error, the key is to turn the original goal into a simultaneous inference problem where the recommended sample size to offset an increased possibility of estimation error is directly related to the number of inferences to be conducted. Given a fixed sample size budget, our result allows us to invert the question to one about the feasible number of treatment arms or partition complexity (e.g. number of decision tree leaves). Using policy trees to learn sub-groups, we evaluate our nominal guarantees on a large publicly-available randomized experiment test data set.
翻译:本文探讨如何确定K臂随机实验的充足样本量,以估计数据驱动子组中的条件反事实期望。子组可由任意特征空间划分算法生成,包括基于相似预测得分对用户进行分箱的规则,或基于学习策略树定义的规则。在明确定义推断目标、最小置信水平及最大误差容限后,核心方法是将原始目标转化为同步推断问题——此时推荐样本量需抵消估算误差增大可能性,该样本量直接与待执行推断次数相关。在固定样本量预算下,我们的结果允许将问题逆向转化为关于可行处理臂数量或划分复杂度(例如决策树叶子节点数)的探讨。通过使用策略树学习子组,我们在大规模公开随机实验测试数据集上验证了标称保证的效力。