The causal dependence in data is often characterized by Directed Acyclic Graphical (DAG) models, widely used in many areas. Causal discovery aims to recover the DAG structure using observational data. This paper focuses on causal discovery with multi-variate count data. We are motivated by real-world web visit data, recording individual user visits to multiple websites. Building a causal diagram can help understand user behavior in transitioning between websites, inspiring operational strategy. A challenge in modeling is user heterogeneity, as users with different backgrounds exhibit varied behaviors. Additionally, social network connections can result in similar behaviors among friends. We introduce personalized Binomial DAG models to address heterogeneity and network dependency between observations, which are common in real-world applications. To learn the proposed DAG model, we develop an algorithm that embeds the network structure into a dimension-reduced covariate, learns each node's neighborhood to reduce the DAG search space, and explores the variance-mean relation to determine the ordering. Simulations show our algorithm outperforms state-of-the-art competitors in heterogeneous data. We demonstrate its practical usefulness on a real-world web visit dataset.
翻译:数据中的因果依赖关系通常由有向无环图(DAG)模型刻画,该模型广泛应用于众多领域。因果发现旨在利用观测数据恢复DAG结构。本文聚焦于多元计数数据的因果发现。我们的研究源于真实世界的网页访问数据,该数据记录了单个用户对多个网站的访问行为。构建因果图有助于理解用户在网站间转换的行为模式,从而启发运营策略。建模过程中的一个挑战是用户异质性——不同背景的用户表现出多样化的行为特征。此外,社交网络连接可能导致朋友间产生相似行为。我们引入个性化二项式DAG模型,以应对实际应用中普遍存在的异质性和观测值间的网络依赖性。为学习所提出的DAG模型,我们开发了一种算法:将网络结构嵌入到降维协变量中,通过学习每个节点的邻域来缩小DAG搜索空间,并利用方差-均值关系确定排序。仿真实验表明,在异质性数据场景下,我们的算法优于现有最优方法。我们通过真实网页访问数据集验证了其实用价值。