We study the problem of graph structure identification, i.e., of recovering the graph of dependencies among time series. We model these time series data as components of the state of linear stochastic networked dynamical systems. We assume partial observability, where the state evolution of only a subset of nodes comprising the network is observed. We devise a new feature vector computed from the observed time series and prove that these features are linearly separable, i.e., there exists a hyperplane that separates the cluster of features associated with connected pairs of nodes from those associated with disconnected pairs. This renders the features amenable to train a variety of classifiers to perform causal inference. In particular, we use these features to train Convolutional Neural Networks (CNNs). The resulting causal inference mechanism outperforms state-of-the-art counterparts w.r.t. sample-complexity. The trained CNNs generalize well over structurally distinct networks (dense or sparse) and noise-level profiles. Remarkably, they also generalize well to real-world networks while trained over a synthetic network (realization of a random graph). Finally, the proposed method consistently reconstructs the graph in a pairwise manner, that is, by deciding if an edge or arrow is present or absent in each pair of nodes, from the corresponding time series of each pair. This fits the framework of large-scale systems, where observation or processing of all nodes in the network is prohibitive.
翻译:我们研究了图结构识别问题,即恢复时间序列间依赖关系的图结构。将这些时间序列数据建模为线性随机网络动力系统状态的分量。我们假设部分可观测性,即仅观测到网络中部分节点的状态演化。我们设计了一种从观测时间序列计算得到的新的特征向量,并证明这些特征是线性可分的,即存在一个超平面能够将连接节点对对应的特征簇与未连接节点对对应的特征簇分开。这使得这些特征适用于训练各种分类器进行因果推断。特别地,我们利用这些特征训练卷积神经网络(CNNs)。所得到的因果推断机制在样本复杂度方面优于最先进的同类方法。训练后的CNN能够很好地泛化到结构不同的网络(稠密或稀疏)和不同噪声水平的场景。值得注意的是,尽管仅在合成网络(随机图的实现)上训练,它们也能很好地泛化到真实世界网络。最后,所提出的方法以逐对方式一致地重建图结构,即通过每对节点对应的时间序列判断该对节点之间是否存在边或箭头。这适用于大规模系统框架,其中对网络中所有节点进行观测或处理是不可行的。