This paper studies the problem of designing an optimal sequence of interventions in a causal graphical model to minimize cumulative regret with respect to the best intervention in hindsight. This is, naturally, posed as a causal bandit problem. The focus is on causal bandits for linear structural equation models (SEMs) and soft interventions. It is assumed that the graph's structure is known and has $N$ nodes. Two linear mechanisms, one soft intervention and one observational, are assumed for each node, giving rise to $2^N$ possible interventions. Majority of the existing causal bandit algorithms assume that at least the interventional distributions of the reward node's parents are fully specified. However, there are $2^N$ such distributions (one corresponding to each intervention), acquiring which becomes prohibitive even in moderate-sized graphs. This paper dispenses with the assumption of knowing these distributions or their marginals. Two algorithms are proposed for the frequentist (UCB-based) and Bayesian (Thompson Sampling-based) settings. The key idea of these algorithms is to avoid directly estimating the $2^N$ reward distributions and instead estimate the parameters that fully specify the SEMs (linear in $N$) and use them to compute the rewards. In both algorithms, under boundedness assumptions on noise and the parameter space, the cumulative regrets scale as $\tilde{\cal O} (d^{L+\frac{1}{2}} \sqrt{NT})$, where $d$ is the graph's maximum degree, and $L$ is the length of its longest causal path. Additionally, a minimax lower of $\Omega(d^{\frac{L}{2}-2}\sqrt{T})$ is presented, which suggests that the achievable and lower bounds conform in their scaling behavior with respect to the horizon $T$ and graph parameters $d$ and $L$.
翻译:本文研究在因果图模型中设计最优干预序列的问题,目标是最小化相对于事后最优干预的累积遗憾。该问题自然被建模为因果bandit问题。本文聚焦于线性结构方程模型(SEM)与软干预下的因果bandit。假设图结构已知且包含$N$个节点。每个节点存在两种线性机制(一种软干预机制与一种观测机制),由此产生$2^N$种可能的干预。现有大多数因果bandit算法假设至少奖励节点父节点的干预分布是完全已知的,然而这$2^N$种分布(每种干预对应一个分布)的获取即使在中规模图上也变得不可行。本文摒弃了已知这些分布或其边际分布的假设,针对频率学派(基于UCB)和贝叶斯学派(基于汤普森采样)两种场景分别提出两种算法。这些算法的核心思想是避免直接估计$2^N$个奖励分布,转而估计完全刻画SEM的参数(参数数量与$N$成线性关系),并利用这些参数计算奖励。在两种算法中,基于噪声与参数空间的有界性假设,累积遗憾量级为$\tilde{\cal O} (d^{L+\frac{1}{2}} \sqrt{NT})$,其中$d$为图的最大度,$L$为最长因果路径长度。此外,本文给出一个下界$\Omega(d^{\frac{L}{2}-2}\sqrt{T})$,表明可实现界与下界在关于时间跨度$T$及图参数$d$、$L$的缩放行为上具有一致性。