Suppose that a cascade (e.g., an epidemic) spreads on an unknown graph, and only the infection times of vertices are observed. What can be learned about the graph from the infection times caused by multiple distinct cascades? Most of the literature on this topic focuses on the task of recovering the entire graph, which requires $\Omega ( \log n)$ cascades for an $n$-vertex bounded degree graph. Here we ask a different question: can the important parts of the graph be estimated from just a few (i.e., constant number) of cascades, even as $n$ grows large? In this work, we focus on identifying super-spreaders (i.e., high-degree vertices) from infection times caused by a Susceptible-Infected process on a graph. Our first main result shows that vertices of degree greater than $n^{3/4}$ can indeed be estimated from a constant number of cascades. Our algorithm for doing so leverages a novel connection between vertex degrees and the second derivative of the cumulative infection curve. Conversely, we show that estimating vertices of degree smaller than $n^{1/2}$ requires at least $\log(n) / \log \log (n)$ cascades. Surprisingly, this matches (up to $\log \log n$ factors) the number of cascades needed to learn the \emph{entire} graph if it is a tree.
翻译:假设一个级联(例如流行病)在未知图上传播,仅观测到节点的感染时间。从多个不同级联的感染时间中能推断出图的哪些信息?现有文献主要关注恢复整个图的任务,对于具有有界度的n节点图,这需要Ω(log n)个级联。本文提出不同问题:能否仅通过少量(即常数个)级联,即使在n很大时,也能估计出图的重要部分?本研究聚焦于从图上的易感-感染过程中产生的感染时间识别超级传播者(即高度数节点)。我们的第一个主要结果表明,度数大于n^{3/4}的节点确实可通过常数个级联进行估计。为此提出的算法利用了节点度数与累积感染曲线二阶导数之间的新型关联。反之,我们证明估计度数小于n^{1/2}的节点至少需要log(n)/log log(n)个级联。令人惊讶的是,这(在对数log n因子范围内)与当图为树时学习《整个》图所需的级联数量相匹配。